> ## Documentation Index
> Fetch the complete documentation index at: https://docs.primeintellect.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Volume-hosted models

> Boot a run from weights already on a volume instead of the cluster model cache

<Note>
  Volume-hosted models are part of the [Volumes](/hosted-training/volumes) closed beta. Hosted SFT is supported today; RL runs are coming.
</Note>

Full fine-tuning runs normally boot from the cluster model cache: a run only dispatches to a cluster that already has the base model cached, and [asking us to cache a new model](/hosted-training/full-finetuning#before-you-start) is a manual step.

A run launched with a [volume](/hosted-training/volumes) can instead read its base model straight from the volume: point `[model].name` at the weights with an absolute `/volume/...` path, and the cache is not involved at all. No cache request, no `prime train models` check — the weights are yours and they are already where the run will read them.

## When this matters

* **Chaining runs on your own fine-tunes.** A run with `[ckpt]` saves its checkpoint to the volume. Point the next run's `[resume].dir` at that checkpoint's step directory and keep training where the last run stopped — no re-upload, no cache step ([below](#chaining-from-a-previous-runs-checkpoint)).
* **Weights that are not on Hugging Face.** Private or proprietary checkpoints can be staged onto the volume directly and trained from.
* **Iterating offline.** The model you want is never "not available on this cluster" — if it is on the volume, it is available.

## Quick start

Weights staged on the volume `research`:

```bash theme={null}
prime volumes put research ./my-model /models/glm-5.3-hf
```

The directory must contain the weights in Hugging Face format (`config.json` plus safetensors files).

```toml theme={null}
[ckpt]                                # save the result back to the volume

[model]
name = "/volume/models/glm-5.3-hf"    # absolute /volume/... path

[data]
name = "datasets/my-dataset"          # the dataset can live on the same volume
type = "sft"
seq_len = 4096
```

```bash theme={null}
prime train sft.toml --volume research
```

To continue from an earlier run's checkpoint instead, do **not** point `[model].name` at it — run checkpoints are saved in prime-rl's distributed format, not HF format. Use `[resume].dir` with the checkpoint's step directory on the volume ([below](#chaining-from-a-previous-runs-checkpoint)).

## Rules

| | |
| - | - |
| Run type | **Hosted SFT only** — configs with `[data]` and no `[trainer]`/`[orchestrator]`. RL runs with a `/volume/...` model are rejected for now (see [below](#whats-still-rl-only)). |
| Volume required | The run must carry the volume: `--volume <name>` or `volume = "<name>"` in the config. Without it, dispatch is rejected. |
| Path shape | Absolute mount path starting with `/volume/`, at least one directory after it. Unlike `data.name`, volume-relative paths are not accepted here. |
| Segment characters | Each path segment starts with a letter or digit, then letters, digits, `.`, `-`, or `_` only — `.` and `..` are rejected. |
| Weight format | A directory of Hugging Face-format weights (`config.json` plus safetensors files). |
| Cache | Not consulted. No model-cache mount and no cache gate applies. |

## How dispatch changes

Without the volume-model path, dispatch requires a cluster that has the base model cached and fits the run's GPU requirements. With `/volume/...` weights, only the GPU requirement remains: the run dispatches to the volume's cluster as usual (a volume always pins its runs to the cluster it was created on), and nothing checks the cache first.

Inside the trainer pod, the volume is mounted read-only at `/volume`. If the trainer needs to convert the weights (for example HF-to-prime conversion for MoE models), the converted copy is written to the run's own scratch space — never back onto the volume. Run outputs still land under `runs/<runId>/` on the volume, exactly like any other volume run.

## Chaining from a previous run's checkpoint

Checkpoints a run saves under `runs/<runId>/` on the volume are in prime-rl's distributed checkpoint format — not HF-format weights — so they cannot be used as `[model].name`. To continue one, point `[resume].dir` at the checkpoint's step directory on the volume:

```toml theme={null}
[ckpt]

[model]
name = "/volume/models/glm-5.3-hf"  # base weights: still HF format

[data]
name = "datasets/my-dataset"
type = "sft"
seq_len = 4096

[resume]
dir = "/volume/runs/<previousRunId>/outputs/fft-<previousRunId>/checkpoints/step_100"
```

The trainer restores weights, optimizer, dataloader state and the step counter from the checkpoint, then trains the remaining steps: `max_steps` counts from the start of the original training, so set it above the checkpoint's step (here `> 100`) for the fork to train.

Requirements and behavior:

* Resume from a **completed** checkpoint: a run (or step) that finished writing. A mid-write `step_N` directory can fail to load.
* The source run's own retention policy applies to its checkpoints — keep the source checkpoint for as long as the new run may still start or restart against it.
* On trainer pod restarts the run resumes from the same external checkpoint again — not from the newer checkpoints it has saved itself. To fork from the latest state instead, update `resume.dir` and start a new run.

## What's still RL-only

RL runs (`[trainer]`/`[orchestrator]` configs) still require a cached HF model id: their `[trainer].model` is loaded in the trainer, the inference engine, and the orchestrator, so the training chart must expose `/volume` to all three before `/volume/...` model paths can be enabled for RL. When the chart publishes that support, this page will be updated and the restriction dropped.

## See also

* [Volumes](/hosted-training/volumes) — creating, growing, and SSH-ing into volumes; the compact version of this feature's rules lives there too.
* [Hosted SFT](/hosted-training/hosted-sft) — the walkthrough this feature builds on, including dataset staging on a volume.
* [Full fine-tuning](/hosted-training/full-finetuning) — the cached-model flow this replaces for volume runs.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.