eval entrypoint:
.toml files for configuration:
uv run eval @ config.toml --dry-run. To run the evaluation, use uv run eval @ config.toml.
Use dotted arguments to set values using the CLI, e.g. --sampling.temperature 0.5. CLI arguments overwrite toml arguments when both are present.
The output from evaluations are written into outputs/<env>--<model>--<harness>/<uuid>/ by default, where <env> is the taskset, prefixed by the paired env id when --env.id sets one (use output_dir to overwrite the folder). The folder contains the used config.toml, all the episodes in traces.jsonl, as well as logs of the run and workers in eval.log.
Common config values
model— the model id to evaluate, e.g.nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55Bsampling— generation params passed to the model, e.g.sampling.temperatureenv.taskset.id— pick the taskset (or the positionaleval <taskset-id>)env.agent.harness.id— pick the agent’s harness ([env.agent.harness]in TOML)num_tasks— how many tasks to evaluate. Not setting a value means all tasks; an infinite taskset (a procedural generator, e.g.wordle-v1) requires itnum_rollouts— rollouts per taskverbose— log at debug instead of infoshuffle— samples the task order (fixed seed); an error on an infinite taskset
Resuming evaluations
--resume <output-dir> re-runs only the rollouts a previous run left missing or errored, appending to that run’s own traces.jsonl. It reloads the run’s saved config.toml verbatim, so it takes no other arguments. Good rollouts are kept, while errored ones are dropped and redone.
Disabling tools
Almost every harness comes with adisabled_tools list, which can be used to disable one or multiple tools:
Skills
Harnesses whose program supports SKILL.md skills natively (e.g. Claude Code, Codex) take askills list of local skill folders, each uploaded into the program’s skill discovery directory in the agent’s runtime as <skills dir>/<folder name>:
skills on a harness without native skill support fails up front.