Answers
How do I set up kohya_ss or LLaMA-Factory on RunPod or Vast without dependency conflicts every time?
Get the training run working once, then pack it: a nest keeps the training tool, the exact package versions that resolved on that machine, the base model and how the run was started. On the next machine a restore reinstalls exactly those versions and runs the training once, and the run only counts when the adapter file it was packed to produce appears. Your training images don't travel; you put them back.
Why does a fresh install on RunPod give me conflicted versions every time?
Because pip install -r requirements.txt re-decides the versions each time
you run it, against whatever the package index holds that day. Users of both tools ask for
the same thing — "python requirements.txt doesn't install all the right versions …
please lock them down" — and the honest fix is to keep the list that worked, not the
file that produced it. A fresh
resolve is a different environment explains why.
Renest does that for a run that already finished: it records the versions that were actually installed when the training worked, and on the next machine installs exactly that list with uv, with no new resolution step.
How do I save a kohya_ss or LLaMA-Factory setup that works?
Pack it right after a training run that finished. You describe that run in a small JSON
file — the working directory, the command line and the environment
({"cwd": …, "argv": [...], "env": {…}}) — and Renest reads
the recipe out of it: for kohya_ss the command line itself, for LLaMA-Factory the config file
the command names.
$ renest pack --dir ./run --framework kohya --run-record run.json --out ./nests --dest hosted $ renest pack --dir ./run --framework llamafactory --run-record run.json --out ./nests --dest hosted
What goes in: the training tool's folder, the config files the run read, the base model
it trained from, and the dependency lock. What stays out: the checkpoints and adapters the
run produced, and your dataset. They aren't dropped silently — each is recorded in the nest
as a spot that pointed at your own files. --dest hosted uploads the nest to your
drive as it packs, so it isn't left only on a machine you are about to destroy.
Capture a run has the details.
When the dependency list can't be installed elsewhere as recorded — kohya_ss's own
pip install -e . line, packages from conda's base environment — the pack still
finishes, names those packages and the fix, and marks the nest as not restorable on another
machine. See why pip freeze isn't a
backup.
Do I need renest watch?
Run the training once under renest watch before packing. A trainer exits when
it's done, so by packing time there's no process left to ask which system libraries it
loaded. watch records that while the run happens, inside the folder you will
pack, and changes nothing about the run — same arguments, same output, same exit code:
$ renest watch -- accelerate launch train_network.py --config config.toml
Without it, packing falls back to what the installed packages declare, which is wider than what the run used and can only ever warn.
How do I get it running on another RunPod or vast.ai machine?
In the web console, open the nest and choose Restore; the page gives you a command with a restore code. On the new machine:
$ uv tool install renest $ renest restore --grant grant.json --dir ./run
Before any file moves, the restore checks the machine against the one the nest was packed on: a GPU generation the recorded build has no code for, a driver too old for its CUDA build, or a CPU without avx2 stops it there. Missing system libraries are named, with the command that installs them, but don't stop it. It also asks whether the recorded package list can be installed here, before the large downloads.
Then every file is checked against its recorded checksum, the packages are reinstalled
from the lock, and the training command runs once. Its output goes to
.renest/evidence/<run>/entrypoint.log inside the folder you restored to;
when the run fails, that log is the real diagnosis and the restore points you to it. A run
that dies on a missing system library such as libGL.so.1 is reported as that,
with the install command, not as a broken environment. More in
Restore anywhere.
Where did my training data go?
Your dataset is never packed into a nest, and a restore says so as soon as the files have landed: it lists each spot in the recipe that pointed at your own files — the training data it read, where results went — so you know what to put back. If the run then fails because the data isn't there (kohya prints a "no data found" kind of message), the restore says exactly that: put the data back where the recipe expects it and run the same restore command again. Files already checked are kept.
When does a restored training setup count as working?
When the run produces the adapter it was packed to produce. An exit code alone isn't trusted: kohya can print an error and exit 0 when it finds no data, so when a nest names the expected adapter file, the restore requires that file to appear and fails if it doesn't.
Our own fine-tuning restores, failures included, are published on the Proof page, listed separately from image workflows; there a fine-tuning restore counts only when it produces a non-empty trained adapter.
When Renest isn't the answer
- You have never had the training run working. Renest only carries a run that already finished. It doesn't fix the first install, resolve dependency conflicts or recommend a template; the tool's own install guide or a community image is the place to start.
- You train with something other than kohya_ss or LLaMA-Factory — for example diffusion-pipe, musubi, ai-toolkit or Axolotl. Those aren't supported.
- Your new card is a GPU generation the recorded PyTorch build doesn't support. Renest refuses that pairing before downloading; it doesn't swap your PyTorch version. Get the run working on that generation once and pack again.
Related: Keep a setup between cloud sessions · What a backup has to include · The official base image
These docs describe renest 0.1.15, the latest release.