renest 0.1.15

Answers

My network volume has my custom nodes and models. Why does ComfyUI still break on a new pod?

Because a volume keeps files, not a working environment. A custom node needs three things to load: its folder, the Python packages it installed, and the system libraries those packages link against. The folder survives on the volume; a virtual environment built on one machine often fails on the next, and system libraries live in the container image, not on the volume.

Whenever I start a new pod and attach my volume, custom nodes throw red errors. Why?

Because one of the three things a node needs is missing on the new pod, even though its folder is right there on the volume:

  • The node's folder in custom_nodes. This one is on the volume, so it survives.
  • The Python packages it installed. If they went into the system Python on the container disk, they were erased with the old pod. If they are in a virtual environment on the volume, they may not work on the new pod (see below).
  • The system libraries those packages load, such as libGL.so.1. These come from the container image. A volume never holds them.

A red node in the graph is ComfyUI saying the node's class isn't registered, and the reason is in the startup log, not in the graph.

ComfyUI custom nodes say IMPORT FAILED after a server stop. What's missing?

Read the lines right after IMPORT FAILED (or Cannot import … module for custom node) in ComfyUI's startup log. They name the cause:

  • ModuleNotFoundError: No module named … means a Python package is missing. Install the node's requirements into the Python that runs ComfyUI.
  • … cannot open shared object file means a system library is missing. Installing the node again won't fix it; install the library with the system package manager (for example apt-get install), or use an image that already has it.

The nodes missing or import failed page goes through the common cases.

Should I just keep the venv on /workspace?

It helps when the next pod is close to the last one: same image, same Python, a similar GPU. A virtual environment is built for one machine. It points at the Python it was created with, by path, and the PyTorch inside it was built for a particular CUDA version and range of GPU generations. Attach the same volume to a pod with a different image or a newer GPU and that environment can fail to start, or start and then fail with errors such as no kernel image is available for execution on the device; see errors after changing GPU. It also does nothing for system libraries.

Why do I have to pip install the same things again on every new pod?

Because the packages lived on a disk that didn't come with you. Reinstalling from each node's requirements.txt works, but it installs whatever versions are newest today, which may not be the versions your workflow last worked with. To get the same environment back, you need the exact versions that worked, written down when they worked; why pip freeze isn't a backup covers where a plain freeze falls short.

What does a nest bring back that the volume doesn't?

A nest is packed from a run that worked. It records:

  • the files: model weights, the custom nodes' source pinned to their commits, the workflow, each with a recorded checksum;
  • the exact Python package versions that run used, as a lock that can be installed again on another machine;
  • which system libraries the run actually loaded.

On restore, the packages are installed again from that lock for the new machine, instead of copying an environment built for the old one. Before the big download starts, the restore names any of the recorded system libraries this machine lacks, with the command to install them. That is a warning; the restore carries on. After it starts ComfyUI, it reads the log and names any custom node that failed to load because a system library is missing.

The same check refuses a machine before the download when the GPU generation is one the nest's PyTorch build has no code for, when the driver is too old for the nest's CUDA build, or when the CPU lacks avx2. Renest does not swap in a different PyTorch for a new GPU: to run on a newer generation, get the setup working there once and pack again.

How do I move a volume setup into a nest?

On a pod where the workflow still produces images, install the tool and an access key (see the RunPod guide), export the workflow in API format, and pack:

Shell
$ renest pack --dir /workspace --workflow workflow-api.json \
    --out /workspace/nests --dest hosted

On the next pod, open the nest in the web console, choose Restore, and paste the command it gives you. It writes the restore code to grant.json:

Shell
$ renest restore --grant grant.json --dir ./run
$ renest start --dir ./run --listen 0.0.0.0

The restore downloads every file and checks it, installs the locked packages, starts ComfyUI and runs the workflow once, then reports which steps passed. Files coming back checked and ComfyUI working are counted separately; see the Proof page for our own results. Why the two differ is told in Every file verified, environment still dead. If a step stops, see Troubleshooting.

When Renest isn't the answer

If every new pod uses the same image and the same GPU type, and the venv on your volume keeps working, you don't need to change anything. If only one or two system libraries are missing, adding them to your own image or to the pod's start command is the smaller fix. If you need a fixed environment for serverless or production deployment, build a Docker image. See volume, image, snapshot or nest for what each one keeps.

Open format

The docs describe a format you own.

Everything here is written against the open nest format. Every nest ships a plain restore.sh that brings back its files and dependencies with zero Renest code.