# Whole files, whole fingerprints.

> Renest deduplicates complete files: identical bytes, identified by one sha256 fingerprint, are stored once. A file with even one byte changed counts as a new file and is stored and transferred in full.

Canonical page: https://renest.ai/docs-whole-file-dedup.html

<h2 id="what">Renest stores identical files once</h2>

<p>Every model file, source archive, lock file and workflow in a nest is stored under the
sha256 of its complete contents (see <a href="docs-content-addressed-files.html">Files
are known by their bytes</a>). Identical bytes therefore land on the same object:</p>

<ul>
  <li><strong>In your own bucket</strong>, a file that is already there under its
  fingerprint, at the expected size, is not uploaded again.</li>
  <li><strong>On the hosted drive</strong>, files you have stored before are not uploaded
  again. For a larger file the drive already holds but you haven't stored, your tool can
  prove it has the same bytes by answering a spot check on byte ranges of your local
  copy; if the check doesn't pass, the file is uploaded in full.</li>
  <li><strong>On the way back</strong>, a file that appears at several paths in one nest is
  downloaded once and copied, and a rebuild that is run again on the same machine does not
  re-download files it already verified.</li>
</ul>

<p>For the way AIGC work accumulates, this covers a lot: the multi-gigabyte base
checkpoints are the bulk of the bytes, and they are exactly the files that repeat
unchanged from nest to nest. Seal ten variations of a workflow into the same bucket and
the base model is stored once, not ten times.</p>

<h2 id="edge">Changing one byte of a file makes it a new file</h2>

<p>The unit of "same" is the whole file. <strong>Change one byte inside a 6 GB model and,
as far as deduplication is concerned, you have a brand-new 6 GB file</strong> — stored in
full, transferred in full. There is no diffing against the previous version, no
splitting files into chunks and reusing the unchanged ones.</p>

<p>This bites in a specific case: a file that is <em>semantically</em> the same but not
byte-identical. Re-save a checkpoint with different embedded metadata and the tensors
inside may be untouched while the file's fingerprint — computed over header and all —
changes, so both versions are stored. Worth keeping in proportion: the common cases still
deduplicate. The same published file downloaded by everyone is byte-identical everywhere,
and renaming a file changes nothing about its bytes. The gap is the narrower set of files
that were re-written, not merely re-named.</p>

<p>We'd rather you learn this trade-off from this page than from a transfer that's bigger
than you expected. It is a known limitation and a deliberate one.</p>

<h2 id="why">The whole-file sha256 is the identity layer, so chunk-level deduplication is not built</h2>

<p>Chunk-level or tensor-level deduplication would genuinely save space for re-written
files. The reason it isn't here is that the obvious implementation breaks something we
won't break.</p>

<p>In Renest, the whole-file sha256 is not just a storage key. It is the <strong>identity
layer</strong>: it's what the manifest records, what a rebuild verifies against, what the
<a href="docs-escape-hatch.html">escape-hatch script</a> checks with ordinary command-line tools, and what lines up with the fingerprints Hugging Face and Civitai
already publish. A scheme that replaced that fingerprint with per-chunk fingerprints or a
different algorithm would optimise storage by giving up verification with ordinary tools
and ecosystem compatibility. Storage economics are our problem; verifiability is
yours.</p>

<h2 id="later">Any finer-grained deduplication would have to sit beneath the whole-file fingerprint</h2>

<p>If finer-grained deduplication is ever added, it has to be <strong>a second key
underneath, never a different fingerprint on top</strong>. The whole-file sha256 would
stay exactly what it is: the identity every manifest records and every rebuild verifies.
A finer-grained key (for instance, one computed over tensor payloads independently of a
mutable file header) would live inside storage only, with manifests, verification and the
escape hatch unchanged.</p>

<p>No date is attached to that, and this page is not a promise that it's coming. It states
the constraint any future version has to satisfy: <strong>the fingerprint you verify with
is not negotiable, and savings have to be found beneath it, not instead of it.</strong>
Until then, deduplication stays whole-file.</p>
