evaluate
View MarkdownBenchmark evaluation on standard tasks, or over an Environment’s held-out tasks.
| Field | Value |
|---|---|
| Reference | nodus/evaluate |
| Image | nodus/evaluate:2.0.0 |
| Task | Evaluate |
| Modes | Evaluate |
| Launcher | Torchrun |
| Requires | model yes, data no, environment no |
| Checkpoints | /nodus/state (HFTrainer) |
Parameters
Section titled “Parameters”Set under spec.parameters; the server validates them against this schema, and unset parameters take the runtime default.
| Parameter | Type | Default | Allowed | Description |
|---|---|---|---|---|
batchSize |
integer | 8 |
1 to 1024 | Examples per forward pass |
evaluationBatchSize |
integer | none | 1 to 1024 | Prompts generated together during Environment baseline and final evaluation |
limit |
integer | none | 1 to 1000000 | Samples per task (unset scores the whole task) |
maxCompletionLength |
integer | none | 8 to 32768 | Environment evaluation: tokens per completion |
numFewshot |
integer | none | 0 to 64 | Few-shot examples per prompt |
seed |
integer | 0 |
0 to 2147483647 | Seed for sampling and few-shot selection |
taskFilter |
object | none | at most 8 keys | Keep only tasks whose metadata has these values (such as family on reasoning-gym), before trainTasks and heldOutTasks count them |
tasks |
array of string | none | at most 64 items | Benchmark task names, such as arc_easy or hellaswag, whose datasets the run downloads from the Hugging Face Hub before it scores them offline; set either tasks or spec.environment, which scores an Environment’s held-out tasks instead |
Outputs
Section titled “Outputs”| Output | Path |
|---|---|
outputs |
/nodus/outputs |
Every run also writes results.json, provenance.json and a sha256 manifest.json. Each file under /nodus/outputs is an output of the TrainingJob, named by its path: download one with nodus cp tj/<name>:outputs/results.json ./results.json.
Default resources: 1 × RTX-4090 or L40S or RTX-A6000 or A100-80G.
Estimates
Section titled “Estimates”The dry-run estimate multiplies these by the run’s steps or tasks; the nightly runs re-measure them.
| Accelerator | Seconds | Per |
|---|---|---|
| A100-80G | 0.17 | secondsPerTask |
| H100 | 0.11 | secondsPerTask |
| L40S | 0.25 | secondsPerTask |
| RTX-4090 | 0.3 | secondsPerTask |
| RTX-A6000 | 0.33 | secondsPerTask |
Example
Section titled “Example”apiVersion: nodus.dev/v1beta1kind: TrainingJobmetadata: {name: my-evaluate}spec: runtime: nodus/evaluate mode: Evaluate model: {uri: "hf://Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca"} parameters: {tasks: [arc_easy], limit: 200} maxCostUSD: "5.00"