> ## Documentation Index
> Fetch the complete documentation index at: https://docs.primeintellect.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Hosted SFT (Beta)

> Supervised fine-tuning on Hosted Training: volumes, dataset uploads over SSH, and prime train dispatch

<Note>
  Hosted SFT is in **closed beta** — Prime enables it per account. Contact Prime support if dispatch is refused.
</Note>

Run `prime-rl` supervised fine-tuning on Prime's hosted GPU clusters: create a volume, copy your dataset onto it over SSH, dispatch. No `kubectl`, no cluster credentials — the trainer reads the dataset from the volume.

This walkthrough fine-tunes a small Qwen3 model to reverse text ([`reverse-text`](https://github.com/PrimeIntellect-ai/prime-rl/tree/main/examples/basic/reverse-text)) on one GPU. For SFT on your own infrastructure, see [Training](/prime-rl/training).

## Prerequisites

```bash theme={null}
uv tool install prime
prime login
```

**Requires `prime` CLI v0.9.0 or later.**

* A Prime account with hosted training access — Prime-side enablement; the platform deployment needs the hosted SFT path switched on.
* A model **cached on the cluster that owns your volume** — shared-LoRA models are not valid here. Check with:

```bash theme={null}
prime train models --fft-only
```

* A Hugging Face dataset you can download with `hf download` — with `prompt`/`completion` or `messages` columns ([Training § Dataset Format](/prime-rl/training#dataset-format)).

## 1. Create a volume

```bash theme={null}
prime volumes create research --size 100Gi
prime volumes list   # wait for RUNNING (~10 s) before the next step
```

<Note>
  Size for what you keep: datasets plus weights + optimizer state per checkpoint — the 0.6B smoke run wrote \~7 GB; a 100B-class `[ckpt]` needs terabytes. Grow with `prime volumes resize <name> --size <size>`.
</Note>

## 2. Put the dataset on the volume

SSH into the volume and download the dataset onto it. The session runs next to the volume in the cluster, so the download uses the cluster's network — not your laptop:

```bash theme={null}
prime volumes ssh research --read-write
```

Inside the session:

```bash theme={null}
hf download willcb/R1-reverse-wikipedia-paragraphs-v1-1000 \
  --repo-type dataset \
  --local-dir /volume/datasets/reverse-text
exit
```

The trainer sees the volume at `/volume` — the same path you see in the SSH session. Local files work too: `prime volumes put <name> ./my-data /volume/datasets/my-data`. In the config, `data.name` accepts either form: a path relative to the volume (`datasets/my-data`) or the absolute mount path (`/volume/datasets/my-data`) — the platform resolves both to the same location.

## 3. Write the config

Save this as `sft.toml` — `data.name` points at the dataset on the volume (relative to the volume root; absolute `/volume/...` paths work too):

```toml theme={null}
max_steps = 10

[ckpt]               # checkpoint at the end of training, saved to the volume

[model]
name = "PrimeIntellect/Qwen3-0.6B-Reverse-Text-SFT"  # must be in the cluster model cache

[data]
name = "datasets/reverse-text"  # volume-relative; the platform resolves it to /volume/datasets/reverse-text
type = "sft"
seq_len = 2048        # add splits = ["train"] for multi-split datasets
batch_size = 8

[optim]
lr = 2e-5

[deployment]
type = "single_node"
num_train_gpus = 1
gpus_per_node = 1

[renderer]
name = "prime-qwen3"  # must match the model family
```

<Note>
  Hosted SFT is trainer-only: `[eval]`, `[inference]`, and `[weight_broadcast]` blocks are rejected at dispatch, and `output_dir` is injected by the platform — omit both. The walkthrough model is a cached debug sibling already fine-tuned on this dataset — treat the loss curve as a smoke test of the hosted path.
</Note>

## 4. Dispatch

```bash theme={null}
prime train sft.toml --volume research
```

```text theme={null}
Creating Hosted Training run...
Dispatched hosted run kqeggj5dl7k85bc0mk1i4y5i

Monitor run at:
  https://app.primeintellect.ai/dashboard/training/kqeggj5dl7k85bc0mk1i4y5i
```

Use **the run ID your dispatch printed** (`<runId>`) below. The run moves `PENDING` → `CREATING` → `RUNNING` → `COMPLETED` — expect roughly three minutes on one H200 for this walkthrough.

## 5. Watch it run

```bash theme={null}
prime train get <runId>
prime train metrics <runId>   # the loss series the dashboard renders
prime train logs <runId>      # the live trainer log
```

```text theme={null}
Run kqeggj5dl7k85bc0mk1i4y5i

  Status: COMPLETED
  Model: PrimeIntellect/Qwen3-0.6B-Reverse-Text-SFT
  Max Steps: 10
  Created: 2026-09-28 20:33
  Started: 2026-09-28 20:34
  Completed: 2026-09-28 20:37
```

```text theme={null}
step: [1, 2, 3, 4, 5, 6, 7, 8, 9]
loss/mean:       [1.108, 1.1788, 0.9996, 0.9468, 1.1179, 1.044, 1.0353, 1.0038, 0.9944]
loss/perplexity: [3.028, 3.251, 2.717, 2.577, 3.058, 2.841, 2.816, 2.729, 2.703]
```

```text theme={null}
20:35:38 [INFO] Initializing scheduler with 10 steps (type='constant')
Generating train split: 100%|████████████████████| 1000/1000 [00:00<00:00, 23213.48
examples/s]
20:35:47 [SUCCESS] Step 1 |    8.5s | Loss 1.1080 | Grad. Norm 12.9495 | LR
2.00e-05 | Throughput 0 tokens/s | MFU 0.0% | Peak Mem. 10.7/139.8 GiB (7.7%)
20:35:52 [SUCCESS] Step 3 |    2.4s | Loss 0.9996 | Grad. Norm 7.4464 | LR
2.00e-05 | Throughput 6704 tokens/s | MFU 2.9% | Peak Mem. 10.7/139.8 GiB (7.7%)
```

Ten steps on a 0.6B model with a constant LR will not descend cleanly — the point is a visible, NaN-free curve. Occasionally the last metrics row lands during teardown; terminal status and the trainer's final log lines are the success signal. Teardown is automatic; outputs and the `datasets/` tree stay on the volume.

Fetch your outputs any time:

```bash theme={null}
prime volumes get research /runs/<runId>/outputs/fft-<runId>/checkpoints/step_10/trainer .
```

(Usage accounting shows `0 tokens / $0.00` for SFT — a reporting quirk, not free usage.)

## Troubleshooting

<AccordionGroup>
  <Accordion title="[eval] / [inference] blocks rejected">
    Hosted SFT is trainer-only — `[weight_broadcast]` is rejected for the same reason. Remove the blocks, or run locally with `uv run sft`.

    ```text theme={null}
    Error: hosted SFT runs are trainer-only — [eval], [inference] (online evals) are
    not supported on the dedicated path. Remove them, or run locally with prime-rl's
    SFT launcher.
    ```
  </Accordion>

  <Accordion title="Model not available on this cluster">
    Hosted SFT boots from the cluster model cache. Check `prime train models --fft-only`, or ask Prime support to cache the model.

    ```text theme={null}
    Error: HTTP 400: Model 'PrimeIntellect/Qwen3-0.6B' is not available on this
    cluster. Reach out to Prime support to get it added.
    ```
  </Accordion>

  <Accordion title="data.name requires --volume">
    Real data needs a volume. `data.type = "fake"` is the one no-volume path (synthetic data, connectivity smoke test).

    ```text theme={null}
    Error: HTTP 400: data.name requires `--volume <name>`. Create a volume with
    `prime volumes create <name>`, then rerun `prime train sft.toml --volume <name>`.
    ```
  </Accordion>

  <Accordion title="Dataset missing on the volume">
    The trainer reads `data.name` from the volume mount — a path that is empty or missing fails at startup. Re-run step 2 and check the name matches: the directory you create under `/volume/datasets/<name>` in the SSH session is exactly the `datasets/<name>` you point `data.name` at (relative paths resolve under `/volume`, so `datasets/<name>` and `/volume/datasets/<name>` are the same location).
  </Accordion>

  <Accordion title="Volume ran out of space during download">
    Inside the SSH session, `hf download` fails when the volume is full. Exit, grow the volume, and re-run the download:

    ```bash theme={null}
    prime volumes resize research --size <larger-size>
    prime volumes ssh research --read-write
    ```
  </Accordion>
</AccordionGroup>

## Limitations

* **Requires CLI v0.9.0+** — older versions misroute SFT configs against the RL schema.
* **SFT-shaped datasets** — `prompt`/`completion` or `messages` columns.
* **Cluster-cached models only** — `prime train models --fft-only`; support can add others.
* **Trainer-only** — no online evals; `[eval]`/`[inference]`/`[weight_broadcast]` rejected.
* **Checkpoints live on the volume** — `prime train checkpoints` doesn't list SFT ones; use `prime volumes get`/`ssh`.

## Scaling up

The same flow carries to multi-node — switch `[deployment]` to `multi_node` and add the parallelism knobs. Save as `glm53-sft.toml`:

```toml theme={null}
max_steps = 4

[env_vars]
TRITON_CACHE_DIR = "/tmp/.cache/triton_cache"

[model]
name = "zai-org/GLM-5.3-BF16"
impl = "custom"
attn = "flash_attention_2"
cp = 8
ep = 8
optim_cpu_offload = true

[model.ac]
freq = 1

[model.ac_offloading]
max_inflight_activations = 5

[optim]
type = "sign_sgd"
lr = 5e-5
weight_decay = 0.01

[scheduler]
type = "linear"
warmup_steps = 50
decay_steps = 0

[renderer]
name = "glm-5.3"

[deployment]
type = "multi_node"
num_train_nodes = 8
gpus_per_node = 8

[model.debug]
force_balanced_routing = true  # debug only

[data]
type = "sft"
name = "datasets/intellect-3-sft-10k"  # put it on the volume first (step 2); volume-relative
splits = ["math"]
batch_size = 8
micro_batch_size = 1
seq_len = 4096
```

<Note>
  This config targets 8×8 H200 (64 GPUs) — 4 steps take \~5.5 min on that hardware. If you adapt it: remove `force_balanced_routing` — it forces round-robin expert routing and is debug-only; with `warmup_steps = 50` a 4-step run never leaves warmup, so losses are warmup-dominated rather than convergence; add a `[ckpt]` block if you want checkpoints (this one omits it, so none are saved); with this linear schedule the LR ramps to its peak at step 50 — expect a transient loss spike near the peak before it settles (9.42 at step 55 in a 60-step run). See [Scaling](/prime-rl/scaling) and [Advanced](/prime-rl/advanced) for the knobs.
</Note>

<CardGroup cols={2}>
  <Card title="Full Fine-Tuning" icon="wand-sparkles" href="/hosted-training/full-finetuning">
    RL-style full-parameter training on a dedicated cluster — same CLI, same dashboard.
  </Card>

  <Card title="prime-rl Training" icon="book-open" href="/prime-rl/training">
    The underlying SFT trainer: config schema, dataset format, and local runs.
  </Card>
</CardGroup>
