Volume-hosted models are part of the Volumes closed beta. Hosted SFT is supported today; RL runs are coming.
[model].name at the weights with an absolute /volume/... path, and the cache is not involved at all. No cache request, no prime train models check — the weights are yours and they are already where the run will read them.
When this matters
- Chaining runs on your own fine-tunes. A run with
[ckpt]saves its checkpoint to the volume. Point the next run’s[resume].dirat that checkpoint’s step directory and keep training where the last run stopped — no re-upload, no cache step (below). - Weights that are not on Hugging Face. Private or proprietary checkpoints can be staged onto the volume directly and trained from.
- Iterating offline. The model you want is never “not available on this cluster” — if it is on the volume, it is available.
Quick start
Weights staged on the volumeresearch:
config.json plus safetensors files).
[model].name at it — run checkpoints are saved in prime-rl’s distributed format, not HF format. Use [resume].dir with the checkpoint’s step directory on the volume (below).
Rules
How dispatch changes
Without the volume-model path, dispatch requires a cluster that has the base model cached and fits the run’s GPU requirements. With/volume/... weights, only the GPU requirement remains: the run dispatches to the volume’s cluster as usual (a volume always pins its runs to the cluster it was created on), and nothing checks the cache first.
Inside the trainer pod, the volume is mounted read-only at /volume. If the trainer needs to convert the weights (for example HF-to-prime conversion for MoE models), the converted copy is written to the run’s own scratch space — never back onto the volume. Run outputs still land under runs/<runId>/ on the volume, exactly like any other volume run.
Chaining from a previous run’s checkpoint
Checkpoints a run saves underruns/<runId>/ on the volume are in prime-rl’s distributed checkpoint format — not HF-format weights — so they cannot be used as [model].name. To continue one, point [resume].dir at the checkpoint’s step directory on the volume:
max_steps counts from the start of the original training, so set it above the checkpoint’s step (here > 100) for the fork to train.
Requirements and behavior:
- Resume from a completed checkpoint: a run (or step) that finished writing. A mid-write
step_Ndirectory can fail to load. - The source run’s own retention policy applies to its checkpoints — keep the source checkpoint for as long as the new run may still start or restart against it.
- On trainer pod restarts the run resumes from the same external checkpoint again — not from the newer checkpoints it has saved itself. To fork from the latest state instead, update
resume.dirand start a new run.
What’s still RL-only
RL runs ([trainer]/[orchestrator] configs) still require a cached HF model id: their [trainer].model is loaded in the trainer, the inference engine, and the orchestrator, so the training chart must expose /volume to all three before /volume/... model paths can be enabled for RL. When the chart publishes that support, this page will be updated and the restriction dropped.
See also
- Volumes — creating, growing, and SSH-ing into volumes; the compact version of this feature’s rules lives there too.
- Hosted SFT — the walkthrough this feature builds on, including dataset staging on a volume.
- Full fine-tuning — the cached-model flow this replaces for volume runs.