Skip to main content
v0 is considered deprecated and will be fully removed in a future release.

Getting Started

How do I quickly test my environment?

Install it into your project, then run vf-eval with a small sample:
Add -s to save the results under outputs/evals/.

How do I see what the model is outputting?

If using vf-eval: Run with -s to save results under outputs/evals/<env>--<model>/<run-id>/ (metadata.json and results.jsonl). To browse a run on the dashboard, upload it. Uploads need the environment on the Environments Hub: push it with prime env push if you haven’t, then pass its slug with --env:
If using the Python API (env.generate() / env.evaluate()):

How do I enable debug logging?

Set the VF_LOG_LEVEL environment variable:

Environments

Which environment class should I use?

  • SingleTurnEnv: One prompt, one response (Q&A, classification)
  • MultiTurnEnv: Custom back-and-forth interaction (games, simulations)
  • ToolEnv: Model calls Python functions (search, calculator)
  • StatefulToolEnv: Tools that need per-rollout state (sandbox IDs, sessions)

What does max_turns=-1 mean?

Unlimited turns. The rollout continues until a stop condition is triggered (e.g., model stops calling tools, or a custom condition you define).

How do I add a custom stop condition?

Use the @vf.stop decorator on a method that returns True to end the rollout:

How do I handle tool call errors gracefully?

In ToolEnv, customize error handling:
Non-critical errors are returned to the model as tool responses so it can retry.

Reward Functions

What arguments can my reward function receive?

Reward functions receive any of these via **kwargs:
  • completion - the model’s response
  • answer - ground truth from dataset
  • prompt - the input prompt
  • state - full rollout state
  • parser - the rubric’s parser (if set)
  • task - vf.Task object for taskset-backed environments
  • info - metadata dict from dataset
Just include the ones you need in your function signature.

How do group reward functions work?

Group reward functions receive plural arguments (completions, answers, states) and return a list of floats. They’re detected automatically by parameter names:

Training

How do I use a local vLLM server?

Point the client to your local server:

Which client_type should I use for RL training?

Three options trade off control vs simplicity:
  • openai_chat_completions (MITO) — server-side templating, text only. Standard OpenAI path. The trainer re-tokenizes for training, which can drift across multi-turn rollouts and fragment them into multiple samples.
  • openai_chat_completions_token (TITO) — server-side templating, returns token IDs alongside text. The trainer doesn’t re-tokenize. Use when the server’s chat template is stable across turns.
  • renderer — client-side tokenization via a per-model renderer in the renderers package. Stronger token-preservation in theory: bridge_to_next_turn keeps multi-turn rollouts merged into one sample and survives mid-completion truncation cleanly. Hand-coded renderers exist only for a subset of models and corner cases are still being shaken out.
For production training, use openai_chat_completions_token — it’s the tried-and-tested path. Try renderer if you want the stronger guarantees and your model has a hand-coded renderer. See Inference Client Types for the full breakdown.