Skip to content

Benchmark evaluation on standard tasks, or over an Environment’s held-out tasks.

Field Value
Reference nodus/evaluate
Image nodus/evaluate:2.0.0
Task Evaluate
Modes Evaluate
Launcher Torchrun
Requires model yes, data no, environment no
Checkpoints /nodus/state (HFTrainer)

Set under spec.parameters; the server validates them against this schema, and unset parameters take the runtime default.

Parameter Type Default Allowed Description
batchSize integer 8 1 to 1024 Examples per forward pass
evaluationBatchSize integer none 1 to 1024 Prompts generated together during Environment baseline and final evaluation
limit integer none 1 to 1000000 Samples per task (unset scores the whole task)
maxCompletionLength integer none 8 to 32768 Environment evaluation: tokens per completion
numFewshot integer none 0 to 64 Few-shot examples per prompt
seed integer 0 0 to 2147483647 Seed for sampling and few-shot selection
taskFilter object none at most 8 keys Keep only tasks whose metadata has these values (such as family on reasoning-gym), before trainTasks and heldOutTasks count them
tasks array of string none at most 64 items Benchmark task names, such as arc_easy or hellaswag, whose datasets the run downloads from the Hugging Face Hub before it scores them offline; set either tasks or spec.environment, which scores an Environment’s held-out tasks instead
Output Path
outputs /nodus/outputs

Every run also writes results.json, provenance.json and a sha256 manifest.json. Each file under /nodus/outputs is an output of the TrainingJob, named by its path: download one with nodus cp tj/<name>:outputs/results.json ./results.json.

Default resources: 1 × RTX-4090 or L40S or RTX-A6000 or A100-80G.

The dry-run estimate multiplies these by the run’s steps or tasks; the nightly runs re-measure them.

Accelerator Seconds Per
A100-80G 0.17 secondsPerTask
H100 0.11 secondsPerTask
L40S 0.25 secondsPerTask
RTX-4090 0.3 secondsPerTask
RTX-A6000 0.33 secondsPerTask
apiVersion: nodus.dev/v1beta1
kind: TrainingJob
metadata: {name: my-evaluate}
spec:
runtime: nodus/evaluate
mode: Evaluate
model: {uri: "hf://Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca"}
parameters: {tasks: [arc_easy], limit: 200}
maxCostUSD: "5.00"