Skip to main content
Volume-hosted models are part of the Volumes closed beta. Hosted SFT is supported today; RL runs are coming.
Full fine-tuning runs normally boot from the cluster model cache: a run only dispatches to a cluster that already has the base model cached, and asking us to cache a new model is a manual step. A run launched with a volume can instead read its base model straight from the volume: point [model].name at the weights with an absolute /volume/... path, and the cache is not involved at all. No cache request, no prime train models check — the weights are yours and they are already where the run will read them.

When this matters

  • Chaining runs on your own fine-tunes. A run with [ckpt] saves its checkpoint to the volume. Point the next run’s [resume].dir at that checkpoint’s step directory and keep training where the last run stopped — no re-upload, no cache step (below).
  • Weights that are not on Hugging Face. Private or proprietary checkpoints can be staged onto the volume directly and trained from.
  • Iterating offline. The model you want is never “not available on this cluster” — if it is on the volume, it is available.

Quick start

Weights staged on the volume research:
The directory must contain the weights in Hugging Face format (config.json plus safetensors files).
To continue from an earlier run’s checkpoint instead, do not point [model].name at it — run checkpoints are saved in prime-rl’s distributed format, not HF format. Use [resume].dir with the checkpoint’s step directory on the volume (below).

Rules

How dispatch changes

Without the volume-model path, dispatch requires a cluster that has the base model cached and fits the run’s GPU requirements. With /volume/... weights, only the GPU requirement remains: the run dispatches to the volume’s cluster as usual (a volume always pins its runs to the cluster it was created on), and nothing checks the cache first. Inside the trainer pod, the volume is mounted read-only at /volume. If the trainer needs to convert the weights (for example HF-to-prime conversion for MoE models), the converted copy is written to the run’s own scratch space — never back onto the volume. Run outputs still land under runs/<runId>/ on the volume, exactly like any other volume run.

Chaining from a previous run’s checkpoint

Checkpoints a run saves under runs/<runId>/ on the volume are in prime-rl’s distributed checkpoint format — not HF-format weights — so they cannot be used as [model].name. To continue one, point [resume].dir at the checkpoint’s step directory on the volume:
The trainer restores weights, optimizer, dataloader state and the step counter from the checkpoint, then trains the remaining steps: max_steps counts from the start of the original training, so set it above the checkpoint’s step (here > 100) for the fork to train. Requirements and behavior:
  • Resume from a completed checkpoint: a run (or step) that finished writing. A mid-write step_N directory can fail to load.
  • The source run’s own retention policy applies to its checkpoints — keep the source checkpoint for as long as the new run may still start or restart against it.
  • On trainer pod restarts the run resumes from the same external checkpoint again — not from the newer checkpoints it has saved itself. To fork from the latest state instead, update resume.dir and start a new run.

What’s still RL-only

RL runs ([trainer]/[orchestrator] configs) still require a cached HF model id: their [trainer].model is loaded in the trainer, the inference engine, and the orchestrator, so the training chart must expose /volume to all three before /volume/... model paths can be enabled for RL. When the chart publishes that support, this page will be updated and the restriction dropped.

See also

  • Volumes — creating, growing, and SSH-ing into volumes; the compact version of this feature’s rules lives there too.
  • Hosted SFT — the walkthrough this feature builds on, including dataset staging on a volume.
  • Full fine-tuning — the cached-model flow this replaces for volume runs.