Guides
If a restore stops.
When renest restore stops, its last lines name the stage, a reason class such as [S0/DISK_INSUFFICIENT], and what to do. A refusal at the machine check means nothing has been downloaded yet, so rent a machine that fits and paste the same command. Machine faults are fixed by a replacement machine, not by downloading again. If you're stuck, run renest support --dir ./run and send us the last lines of the output at support@notify.renest.ai.
Stuck? Paste the last lines of the output to us
Run this on the machine, before you stop or terminate it:
$ renest support --dir ./run
Read what it prints, then send it to
support@notify.renest.ai, or paste it into a
ticket from the Support page in the web console. If
renest support says there are no run records, send the last twenty or so lines
of the restore's own output instead.
A stop names the stage, the reason, and usually the fix
A restore runs its stages in order (S0 checking this machine, S1 downloading, S2 putting files in place, S3 installing dependencies, S4 starting up, S5 test render). When one fails, it prints a line like this and exits with a non-zero code:
[restore] ✗ S0 checking this machine(…s)—— [S0/DISK_INSUFFICIENT] This machine did not pass the checks: …
The part in brackets is the stage and the reason class. The text after it is written to be acted on. If it names one thing to do, do that and run the same command again. A restore carries on where it stopped, and files already on disk that match their checksums are kept. The glossary lists what each stage does.
“Only … GiB free, and this nest needs … GiB”: rent a bigger disk
The machine check compares the free space on the disk holding your --dir
with what the nest needs. It adds room for the Python environment to the nest's own size.
Nothing has been downloaded yet. The message goes on Clear some space or rent a bigger
disk. The exit code is 65.
- Check which disk
--diris on../runis relative to the folder you pasted the command in. On RunPod, only paths under/workspaceare on the volume disk. - Rent a machine with a bigger disk, or restore into a folder on a disk with room, then paste the same command.
- Don't use
--forcehere. The download would fill the disk partway.
If a disk fills up during the restore anyway (another process writing, say), the class is
S2/DISK_FULL (exit code 25). Free up space and run the same command
again.
“This GPU (sm_…) is older / newer than … this nest's torch was built for”: rent a different card
The nest's PyTorch build carries GPU code for a range of card generations. The check
compares this card's generation with that range before anything downloads. The class is
S0/ARCH_UNSUPPORTED (exit code 64). The message names the range, for
example Rent a card at sm_80 or newer. Every file would still come back correctly,
and the app still could not run on this card.
Rent a card in the range the message names, and paste the same restore command. The card the nest was packed on, shown as Packed on in the console, ran it. If the mismatch only shows up once the app runs, the message is Your nest rebuilt correctly, but this GPU is not one the packed PyTorch build has kernels for. The fix is the same: a different card, not a second download.
“This nest needs CUDA …, and this machine's driver only goes up to CUDA …”: rent a newer driver
The host's GPU driver is part of the machine, not of the nest, and the nest's dependency
lock needs a newer one. The class is S0/CUDA_BLOCK (exit code
63). A related message is This machine's GPU driver … is older than the …
that … needs. Rent a machine with a newer driver and paste the same command. Your card
is fine.
“Full nvidia-smi failed here” or “No CUDA GPUs are available”: replace the machine
Some rented GPU machines answer simple questions about their card, then fail when a program actually uses it. The check looks for this in two ways and warns rather than refuses, because the fault sometimes clears on its own:
- Full nvidia-smi failed here (…), while the single-fact queries above still answer — that is the shape of a wedged GPU node.
- Asking this machine for … MiB of video memory failed: …
If you see either one, or the app then fails at start-up with No CUDA GPUs are
available in its log, nothing is wrong with your nest. Stop or terminate that
machine, rent another, and paste the same restore command. The code is not tied to a
machine. Downloading again on the same machine won't help.
“This machine is missing a system library the app needs”: install it, then run again
Some libraries belong to the operating system, so no nest carries them. The check looks
for the ones the working run loaded. Missing ones are reported as a warning before the
download. Missing libraries never block the restore on their own. If they are still missing
at the end, the closing lines say Read this before you call it done. If the app then
fails because of one, the class is SYSLIB_MISSING (exit code 46
at start-up, 55 during the test render). Every message names the libraries and
the fix:
- Surest fix: when the nest records the image it was packed on, the message names it. Start a machine from that image.
- Or install the packages it names, for example
apt-get install -y libgl1, then run the same restore command again. Your files are kept. - A library that belongs to the GPU driver can't be installed as a package. The message says so. Use a machine whose driver is set up, or start the container with GPU access.
The Renest base image already includes the system libraries that image and fine-tuning workflows commonly load at runtime.
“Rebuilt successfully, but this machine isn't identical”: this is a success
This line appears with exit code 0. Every file is correct. The machine
differs from the one the nest was packed on, for example in a package version or a system
library, so a render here may differ slightly from the original. Nothing to fix unless it
also names a missing library (see above).
UNTRUSTED_SOURCE: the nest wants to install from a server nobody recognises
The restore stopped before installing anything (S3/UNTRUSTED_SOURCE, exit
code 37). Installing a dependency runs code the server sends. The message lists
the servers and, for a nest someone handed you, names the sender. Decide whether you trust
that source. If you do, run the same command again with the --trust-host
option the message spells out, plus --trust-sender for a nest someone handed
you.
Unrecognised dependency sources walks through the
decision. A nest someone handed you that wants to run setup commands stops the same way,
with S2/UNTRUSTED_SETUP (exit code 26).
“This restore code has expired or been revoked”: issue a new one
The full message is This restore code has expired or been revoked. Sign a new one from
your drive — your nest is still there. The class is S1/CREDENTIAL_EXPIRED
(exit code 13). Open the nest in your drive, choose Restore,
issue a new code and paste the new command with the same --dir. Everything
already downloaded is reused.
A dropped connection: run the same command again
Network trouble while downloading (S1/NETWORK_INTERRUPTED, exit code
11) or while installing dependencies (S3/UPSTREAM_UNREACHABLE, exit
code 38) is temporary. Run the same command again once the network is back. It
carries on where it stopped. If installing dependencies is slow where you are,
--package-source points it at a closer package index. See
Restore anywhere.
“What it does not have is your data”: put your training data back
For a fine-tuning nest, S4/NEED_USER_DATA (exit code 45) means
the environment is back but the training data the config points at is not. Training data
never travels with a nest. Put it where the message says the config expects it, then run the
same command again.
Exit code 3: no access key, or a refused one
Exit code 3 is a configuration or credential problem. It comes from commands
that talk to your drive with an access key: renest pack --dest hosted,
renest list and renest export. The usual message is:
✗ No access token. Generate one in the web console, then put it in the RENEST_TOKEN environment variable, or write it into [auth] token in ~/.config/renest/config.toml.
Create a key in the web console under Settings → Access keys and set it as the message says. If the message says the key was refused, it has probably been revoked. Create a new one. Restoring with a restore code never needs an access key. See Signing in to your drive.
renest support turns a failed run into text you can read and paste
$ renest support --dir ./run
It reads the records the last restore in that folder left in
.renest/evidence/. It prints one block with the tool's version, facts about the
machine, the timings, the stage and reason it stopped, and the last lines of the app's or
training run's log. Before printing, it replaces things that look like credentials (keys,
tokens, signed-link signatures) with a visible marker, and replaces your home folder path
with ~. It only recognises secrets by their shape, so read the block before
you send it.
It works entirely on the machine. It doesn't go online, upload anything or write a file.
The block goes to standard output, so renest support --dir ./run >
support.txt saves it. --run picks an earlier run by its folder name under
.renest/evidence/. By default it uses the newest one.
Exit codes: the tens digit is the stage
0 is success. 2 is a usage problem and 3 a
configuration or credential problem; both happen before any stage starts. Every other code
has two digits. The tens digit is the stage (S1 to S5 give 1x to
5x, and S0 gives 6x), and the ones digit is the reason within it.
x0 means a failure in that stage that has no more specific class. The ones on
this page:
11network interrupted ·13restore code expired or revoked25disk filled during the restore ·26setup commands from a sender you haven't confirmed37unrecognised install source ·38package sources unreachable45training data missing ·46system library missing at start-up53no GPU code for this card at run time ·55system library missing during the test render63driver too old ·64GPU generation not supported ·65not enough disk
These docs describe renest 0.1.15, the latest release.