# A 23GB video nest, and the exact boundary of 'identical'

> Same GPU model: the regenerated video is literally the same file. Different model — even same architecture: a uniform ~0.94 plateau. And the 0.9814 mystery is solved: it was the CPU vendor.

Published: 2026-07-15 · Field reports
Canonical page: https://renest.ai/field-report-vol-3.html

> **Update, September 2026.** Kept as measured in July 2026; the packed environment is now called a *nest*. The CPU-vendor finding is written up in full in [Same GPU, two images](cpu-vendor-decides-identical-images.html). We no longer quote a "typical" door-to-door time; current results are on the [Proof page](proof.html).

This volume asks the question video creators actually care about: with a heavy image-to-video workflow (23GB nest), where exactly does "identical" end? One rule up front: this round is **measurement, not acceptance** — SSIM was recorded, not graded pass/fail, and every table was generated straight from raw log files with zero hand-transcription.

## Same GPU model: the video is the same file

We packed a 23GB Wan i2v environment on an RTX A5000 in Canada and revived it on an A5000 in Sweden. Door to door: **13.0 minutes**. Every file verified byte-for-byte. Then the environment regenerated the video — and produced **the same byte sequence: the sha256 of the output video matched the original**. Not "looks the same." The same file, through a lossy WEBP encoding chain, across an ocean, on two hosts running *different driver versions* (580.126 / 580.159).

Other numbers: transfer sustained 829 Mbps (four large files in parallel — faster than Vol. 1's single-file 217–867 Mbps), verifying 24.6GB took 116s, and thanks to dedup, a failed earlier run's 23.5GB upload meant this run uploaded only **11 MiB**. Failed runs' bytes aren't wasted.

## Off the same model: a uniform ~0.94 plateau

Revive the same nest on an RTX 4000 Ada (one architecture generation newer): restore byte-for-byte, video regenerates fine, but frames sit at ~0.94 SSIM. Here's the finding that rewrote this report **while we were writing it**: an A10 — the *same* Ampere architecture as the A5000, just a different model, revived across clouds — landed at 0.944, nearly the identical curve. Our draft said "same architecture generation = zero loss." The A10 data overturned that on the spot. **The determinism cliff sits at the GPU model boundary, not the architecture boundary.**

A follow-up run untangled the remaining variables: a Vast A5000 in Alberta — different cloud, driver 565 vs. 580, different vBIOS minor version — rebuilt the RunPod nest and produced **the same byte sequence again**. Changing cloud doesn't break it. Changing driver major version doesn't break it. Changing vBIOS minor doesn't break it. **Only changing the GPU model breaks it.** (With one asterisk: same model is necessary, not sufficient — read on.)

## A self-correction: the "cumulative drift" that wasn't

Our first/middle/last frame samples on the Ada run read 0.9713 / 0.9434 / 0.9381 — monotonically decreasing, and our initial judgment said "temporal error accumulation, confirmed." The full 17-frame curve **refuted that**: frame 0 is anchored high (the conditioning image in the nest is byte-identical), frame 1 drops once to the ~0.94 plateau, then flat within ±0.003. The cross-model difference is a uniform offset, **not a drift that worsens with video length**. Three points forming a monotone sequence was a shape coincidence; collecting full curves exists precisely to catch this illusion.

## The 0.9814 mystery, solved the same day: it's the CPU

Vol. 2 left an open audit item: two Japanese A4000 hosts both producing exactly 0.981439, with drivers as the prime suspect. **That suspicion is formally corrected here: the driver is innocent. The culprit is the CPU vendor.** Every historical rebuild image falls into exactly two byte-equivalence classes — images within a class are sha256-identical across countries and hosts. Driver 580.159 appears in *both* classes, eliminating it. All Intel hosts (three CPU generations × drivers 570/580/595) land in class B; all AMD hosts (two generations × drivers 550/580) land in class A. The SSIM is always *exactly* 0.981439 because class B is one deterministic image, not a cloud of near-identical ones — it's a determinate function, not noise.

Caveats we're keeping: the baseline machine's CPU was never sampled (probes didn't collect it then), so "baseline = AMD" is inferred from the byte class, not observed. And whether the *video* pipeline is CPU-vendor-sensitive is not yet settled. The practical rule for now: pixel-exact reproduction needs **same GPU model and same CPU brand**; a different CPU brand gives a stable, deterministic micro-difference (98.1%), not randomness.

## The failures, in full

Getting the cross-cloud runs took **four consecutive failed rentals** before a success: one instance stuck in `loading` for 600s; one host that never picked up the order for 1200s; one where *our own* polling loop crashed on a single TLS hiccup — half our fault, and now fixed; and one offer snatched by someone else in the minutes between search and order. Cost of the four: $0.028 / $0.055 / $0.042 / $0. The one thing that never failed: the guard's destroy-on-exit path ran all three times it was needed — no failure mode ever breached the spending cap. Whole exam: **$0.859**, and of our own five earlier aborted runs, three were preventable bugs on our side. Two new "files verify, environment is dead" modes also joined the catalog: a Pentium host crashing ComfyUI with `Illegal instruction`, and a 545.23 driver that can't host CUDA 12.4.

## Caliber

Ada and A10 are n=1 each. One resolution/length/setting only (480p, 17 frames, 15 steps, fp8); sequences longer than 17 frames unexplored. Whether the 0.94 plateau or the 0.9814 micro-difference is visible to a human eye: not yet evaluated (the side-by-side evidence is archived for exactly that). Generation-time wobble of ±10% is unattributed. Blackwell GPUs: video untested (images already known not to run). The CPU-vendor effect: which operator in the image pipeline causes it is not yet identified. These boundaries are drawn from small samples — sharp where we measured, open where we didn't.
