Short answer: our restore pre-check refused machines whose GPU driver was "too old", and some of those machines worked fine. The floor had been copied from the GPU vendor's minimum-driver table for installing the CUDA toolkit — but Renest installs PyTorch wheels that carry their own CUDA runtime, which have a lower floor. We corrected it three times on 2026-08-10 and settled on the vendor's family minimum (525.60 for CUDA 12.x), confirmed on a rented card by running real GPU work.
Round one: the right table, the wrong column
The pre-check has a line that says, in effect, this machine's GPU driver is too old — rent a different one. Its number came from the vendor's own compatibility documentation, specifically the minimum driver for installing the CUDA toolkit.
We don't install the toolkit. A restore installs pip wheels that bundle their own CUDA runtime, and the toolkit's floor doesn't apply to that path. The result was a confident, official-looking rejection of machines that could run the nest; drivers in the 550 range were common on the machines we rented. So we lowered the floor to 550.54, backed by one of our own runs where an older driver had failed hard.
Round two: the fix was still wrong
An internal review found the same mistake one layer down. For applications that carry their own runtime, the vendor publishes a different set of minimums per CUDA family — and for 12.x that minimum is 525.60. Our "fixed" 550.54 still turned away every good machine between 525 and 550, including the 535 long-term-support line.
Round three: the number that happened to be right
The same review suspected our CUDA 13 floor of 580 had the same bug. We checked: it is correct — the 13.x family minimum is also 580.
That coincidence made the diagnosis precise. The problem was never "we copied a vendor table"; in a container setup the toolkit column is even the right one. The problem was copying a number without recording which install path it describes. The same digits are a criterion on one path and a superstition on another.
And one more layer: the single failure that had propped up the 550.54 floor — an old driver with a CUDA 12.4 stack — was a container observation, used as evidence for a wheel threshold. We kept it and relabelled it with the path it belongs to. Evidence with the wrong label is worse than none, because it lends confidence.
How we loosened it without guessing
Loosening a floor has an asymmetric cost: too strict rejects people up front, visibly; too loose lets them through to fail halfway, which is much worse. So before adopting 525.60 we rented a card with driver 535.154.05, installed the CUDA 12.8 PyTorch wheel, and — not trusting the library's own "CUDA is available" answer — ran a real matrix multiply on the GPU and compared it with the CPU result. The pass/fail criteria were written down and committed before the run. It passed.
That supports "535 works for this wheel". It does not prove every driver above 525.60 works with every wheel we will ever install, which is why the measurement stays on record next to the number.
What changed
Every driver floor in the pre-check now carries a note saying where it comes from and which install path it describes; the numbers ship in the tool's rules data, so you can read them. The rule we keep: a number with a source is a criterion only if the source measured the path you actually take.
Related: every file verified, environment still dead — the host-side failures the pre-check exists to catch.