Answers
Why do my ComfyUI custom nodes and models get wiped when my RunPod pod stops?
A pod's container disk is erased when the pod stops, so anything installed there goes with it: custom nodes, the Python packages they pulled in, and models saved outside the pod's volume. To keep a setup, either put it on a network volume, which keeps the files in one region of one provider, or pack it as a nest, which keeps the files plus a record of the exact dependency versions, and restore it on whichever machine you rent next.
Why are my Python packages wiped on every stop/start, even though /workspace survived?
Because the packages were installed on the container disk, not on the volume. On a
RunPod pod, only the volume (usually mounted at /workspace) outlives a stop.
The container disk holds the operating system, the system Python and everything
pip install put into it, and it starts from the template's image again each
time. So the custom_nodes folder can still be there while the packages those
nodes import are gone, and the nodes fail to load.
The usual community fix is to build ComfyUI's virtual environment under
/workspace. That works as long as the pod comes back on the same image and a
similar GPU. When it doesn't, see
why ComfyUI still breaks on a new pod
when everything is on a network volume.
Whenever I restart my pod, all my data is lost even if I have a volume. Why?
Usually ComfyUI, or the models, were installed outside the volume's mount point. Some
templates install ComfyUI into a folder on the container disk and only put a few things on
/workspace. Check where your ComfyUI folder really is before you rely on the
volume: anything outside the mount point is gone after a stop.
Two more things to know. A pod's own volume disk survives a stop but is deleted when the pod is terminated. And a stopped pod can only restart on the machine it was on; if that machine's GPUs have been rented by someone else in the meantime, it comes back with zero GPUs. vast.ai works in a similar way: a stopped instance keeps its disk, and is billed for it, but it can only resume when that host's GPU is free again.
How do I pause a cloud GPU and resume it later without losing my ComfyUI setup?
You have three options, and each one costs something different:
- Keep the machine stopped. You pay for its disk while it sits there, and it may not have a GPU free when you come back.
- Keep the files on a network volume and attach it to a new pod next time. You pay for the volume while no pod is running, and it stays in one region of one provider.
- Pack the setup, terminate the machine, and restore it on a new one next time. Nothing on the provider's side keeps billing, and the next machine can be in any region or on another provider. You wait for the files to download again on each restore.
Which is cheapest depends on how often you come back and how big your models are; do the sums for your own use. Keep a setup between cloud sessions compares the three in more detail.
I'll just use a network volume. Is that enough?
If you always come back to the same region of the same provider, a network volume is the simpler tool, and you may not need anything else. It keeps files. It is billed while no pod is using it, it is tied to one region, and the GPU you want is not always available in that region. It also doesn't keep the system libraries that live in the container image, which is the most common reason a setup on a volume still breaks on a new pod.
Why did Download All save the models to my computer instead of the pod?
Because your browser did the download. A download button in the page you have open in
your browser saves to the computer the browser runs on, which is your own machine, not the
pod. To get a model onto the pod, download it from the pod itself: in the pod's terminal
with wget, curl or the Hugging Face command-line tool, or with a
custom node that downloads on the server side. Save it under the folder that survives a stop.
How do I keep a working setup as a nest?
Renest keeps a setup only after it has worked: get your workflow producing images on the pod first. Then pack it. A nest holds the model weights, the custom nodes' source pinned to their commits, the workflow, and a lock of the exact Python package versions that run used. Every file gets a recorded checksum.
- Install the tool on the pod and add an access key; the RunPod guide and the vast.ai guide walk through both.
- Export the workflow in API format (Export (API) in ComfyUI) and pack from the
folder that holds
ComfyUI/:
$ renest pack --dir /workspace --workflow workflow-api.json \ --out /workspace/nests --dest hosted
--dest hosted uploads the nest to your Renest drive as it packs; without it
the nest stays on the pod's disk. Prefer a button? The
Nest this run panel in ComfyUI packs to the
machine's disk, and you upload that folder from the web console.
If the dependency lock can't be reinstalled on another machine (for example, some packages came from conda or were installed from a local folder), the pack still finishes, warns you, and marks the nest as not restorable elsewhere, so you can fix it while the working pod is still running. Keeping your own nests in the Renest drive needs a paid plan; see Pricing. Once the nest is in your drive, you can terminate the pod.
What happens on the next machine?
Rent a Linux machine with an NVIDIA GPU, install the tool, open the nest in the web
console, choose Restore, and paste the command it gives you. It writes your
restore code to grant.json and runs:
$ renest restore --grant grant.json --dir ./run $ renest start --dir ./run --listen 0.0.0.0
Before anything large downloads, the restore checks the machine. It refuses a GPU generation the nest's PyTorch build has no code for, a driver too old for the nest's CUDA build, and a CPU without avx2. It names any system library the original run loaded that this machine lacks, with the command to install it; that is a warning, not a stop. It checks that the dependency lock can be installed here before the model files download. Then it downloads every file and checks it against its checksum, installs the locked packages, starts ComfyUI and runs your workflow once.
Restoring needs no access key: the restore code is the credential, and it works on any machine until it expires, so if a rented box turns out to be broken you can paste the same command on the next one. What the end of a restore prints, and how to reach ComfyUI from your browser, is in After the restore. If it stops, the message names the reason; see Troubleshooting.
Files coming back checked is not the same as ComfyUI working, and we count the two separately; our own results, failures included, are on the Proof page. Want to try the path before packing your own? Take a starter nest.
When Renest isn't the answer
If you never leave one region of one provider and your setup survives restarts on a
volume, keep using the volume. If the only problem is that ComfyUI was installed outside
/workspace, move it onto the volume or pick a template that installs it there.
And Renest can only keep a setup that already works; it does not help you get one working
the first time. See When you don't need
Renest.
These docs describe renest 0.1.15, the latest release.