renest 0.1.15

Guides

If a restore stops.

When renest restore stops, its last lines name the stage, a reason class such as [S0/DISK_INSUFFICIENT], and what to do. A refusal at the machine check means nothing has been downloaded yet, so rent a machine that fits and paste the same command. Machine faults are fixed by a replacement machine, not by downloading again. If you're stuck, run renest support --dir ./run and send us the last lines of the output at support@notify.renest.ai.

Stuck? Paste the last lines of the output to us

Run this on the machine, before you stop or terminate it:

Shell
$ renest support --dir ./run

Read what it prints, then send it to support@notify.renest.ai, or paste it into a ticket from the Support page in the web console. If renest support says there are no run records, send the last twenty or so lines of the restore's own output instead.

A stop names the stage, the reason, and usually the fix

A restore runs its stages in order (S0 checking this machine, S1 downloading, S2 putting files in place, S3 installing dependencies, S4 starting up, S5 test render). When one fails, it prints a line like this and exits with a non-zero code:

[restore] ✗ S0 checking this machine(…s)—— [S0/DISK_INSUFFICIENT] This machine did not pass the checks: …

The part in brackets is the stage and the reason class. The text after it is written to be acted on. If it names one thing to do, do that and run the same command again. A restore carries on where it stopped, and files already on disk that match their checksums are kept. The glossary lists what each stage does.

“Only … GiB free, and this nest needs … GiB”: rent a bigger disk

The machine check compares the free space on the disk holding your --dir with what the nest needs. It adds room for the Python environment to the nest's own size. Nothing has been downloaded yet. The message goes on Clear some space or rent a bigger disk. The exit code is 65.

  • Check which disk --dir is on. ./run is relative to the folder you pasted the command in. On RunPod, only paths under /workspace are on the volume disk.
  • Rent a machine with a bigger disk, or restore into a folder on a disk with room, then paste the same command.
  • Don't use --force here. The download would fill the disk partway.

If a disk fills up during the restore anyway (another process writing, say), the class is S2/DISK_FULL (exit code 25). Free up space and run the same command again.

“This GPU (sm_…) is older / newer than … this nest's torch was built for”: rent a different card

The nest's PyTorch build carries GPU code for a range of card generations. The check compares this card's generation with that range before anything downloads. The class is S0/ARCH_UNSUPPORTED (exit code 64). The message names the range, for example Rent a card at sm_80 or newer. Every file would still come back correctly, and the app still could not run on this card.

Rent a card in the range the message names, and paste the same restore command. The card the nest was packed on, shown as Packed on in the console, ran it. If the mismatch only shows up once the app runs, the message is Your nest rebuilt correctly, but this GPU is not one the packed PyTorch build has kernels for. The fix is the same: a different card, not a second download.

“This nest needs CUDA …, and this machine's driver only goes up to CUDA …”: rent a newer driver

The host's GPU driver is part of the machine, not of the nest, and the nest's dependency lock needs a newer one. The class is S0/CUDA_BLOCK (exit code 63). A related message is This machine's GPU driver … is older than the … that … needs. Rent a machine with a newer driver and paste the same command. Your card is fine.

“Full nvidia-smi failed here” or “No CUDA GPUs are available”: replace the machine

Some rented GPU machines answer simple questions about their card, then fail when a program actually uses it. The check looks for this in two ways and warns rather than refuses, because the fault sometimes clears on its own:

  • Full nvidia-smi failed here (…), while the single-fact queries above still answer — that is the shape of a wedged GPU node.
  • Asking this machine for … MiB of video memory failed: …

If you see either one, or the app then fails at start-up with No CUDA GPUs are available in its log, nothing is wrong with your nest. Stop or terminate that machine, rent another, and paste the same restore command. The code is not tied to a machine. Downloading again on the same machine won't help.

“This machine is missing a system library the app needs”: install it, then run again

Some libraries belong to the operating system, so no nest carries them. The check looks for the ones the working run loaded. Missing ones are reported as a warning before the download. Missing libraries never block the restore on their own. If they are still missing at the end, the closing lines say Read this before you call it done. If the app then fails because of one, the class is SYSLIB_MISSING (exit code 46 at start-up, 55 during the test render). Every message names the libraries and the fix:

  • Surest fix: when the nest records the image it was packed on, the message names it. Start a machine from that image.
  • Or install the packages it names, for example apt-get install -y libgl1, then run the same restore command again. Your files are kept.
  • A library that belongs to the GPU driver can't be installed as a package. The message says so. Use a machine whose driver is set up, or start the container with GPU access.

The Renest base image already includes the system libraries that image and fine-tuning workflows commonly load at runtime.

“Rebuilt successfully, but this machine isn't identical”: this is a success

This line appears with exit code 0. Every file is correct. The machine differs from the one the nest was packed on, for example in a package version or a system library, so a render here may differ slightly from the original. Nothing to fix unless it also names a missing library (see above).

UNTRUSTED_SOURCE: the nest wants to install from a server nobody recognises

The restore stopped before installing anything (S3/UNTRUSTED_SOURCE, exit code 37). Installing a dependency runs code the server sends. The message lists the servers and, for a nest someone handed you, names the sender. Decide whether you trust that source. If you do, run the same command again with the --trust-host option the message spells out, plus --trust-sender for a nest someone handed you. Unrecognised dependency sources walks through the decision. A nest someone handed you that wants to run setup commands stops the same way, with S2/UNTRUSTED_SETUP (exit code 26).

“This restore code has expired or been revoked”: issue a new one

The full message is This restore code has expired or been revoked. Sign a new one from your drive — your nest is still there. The class is S1/CREDENTIAL_EXPIRED (exit code 13). Open the nest in your drive, choose Restore, issue a new code and paste the new command with the same --dir. Everything already downloaded is reused.

A dropped connection: run the same command again

Network trouble while downloading (S1/NETWORK_INTERRUPTED, exit code 11) or while installing dependencies (S3/UPSTREAM_UNREACHABLE, exit code 38) is temporary. Run the same command again once the network is back. It carries on where it stopped. If installing dependencies is slow where you are, --package-source points it at a closer package index. See Restore anywhere.

“What it does not have is your data”: put your training data back

For a fine-tuning nest, S4/NEED_USER_DATA (exit code 45) means the environment is back but the training data the config points at is not. Training data never travels with a nest. Put it where the message says the config expects it, then run the same command again.

Exit code 3: no access key, or a refused one

Exit code 3 is a configuration or credential problem. It comes from commands that talk to your drive with an access key: renest pack --dest hosted, renest list and renest export. The usual message is:

✗ No access token. Generate one in the web console, then put it in the RENEST_TOKEN environment variable, or write it into [auth] token in ~/.config/renest/config.toml.

Create a key in the web console under Settings → Access keys and set it as the message says. If the message says the key was refused, it has probably been revoked. Create a new one. Restoring with a restore code never needs an access key. See Signing in to your drive.

renest support turns a failed run into text you can read and paste

Shell
$ renest support --dir ./run

It reads the records the last restore in that folder left in .renest/evidence/. It prints one block with the tool's version, facts about the machine, the timings, the stage and reason it stopped, and the last lines of the app's or training run's log. Before printing, it replaces things that look like credentials (keys, tokens, signed-link signatures) with a visible marker, and replaces your home folder path with ~. It only recognises secrets by their shape, so read the block before you send it.

It works entirely on the machine. It doesn't go online, upload anything or write a file. The block goes to standard output, so renest support --dir ./run > support.txt saves it. --run picks an earlier run by its folder name under .renest/evidence/. By default it uses the newest one.

Exit codes: the tens digit is the stage

0 is success. 2 is a usage problem and 3 a configuration or credential problem; both happen before any stage starts. Every other code has two digits. The tens digit is the stage (S1 to S5 give 1x to 5x, and S0 gives 6x), and the ones digit is the reason within it. x0 means a failure in that stage that has no more specific class. The ones on this page:

  • 11 network interrupted · 13 restore code expired or revoked
  • 25 disk filled during the restore · 26 setup commands from a sender you haven't confirmed
  • 37 unrecognised install source · 38 package sources unreachable
  • 45 training data missing · 46 system library missing at start-up
  • 53 no GPU code for this card at run time · 55 system library missing during the test render
  • 63 driver too old · 64 GPU generation not supported · 65 not enough disk

These docs describe renest 0.1.15, the latest release.

Open format

The docs describe a format you own.

Everything here is written against the open nest format. Every nest ships a plain restore.sh that brings back its files and dependencies with zero Renest code.