# evaluate

> Benchmark evaluation on standard tasks, or over an Environment's held-out tasks

Source: https://nodus-platform-site.pages.dev/docs/reference/runtimes/evaluate/
Build revision: 211ad9f836655b1c3a2668c4693e442471f28614

<!-- Generated by tools/catalogdocs from catalog/objects. Do not edit: run make gen. -->

Beta

TrainingJob and TrainingRuntime are Beta: fields can still change before GA.

Benchmark evaluation on standard tasks, or over an Environment’s held-out tasks.

|Field|Value|
|-|-|
|Reference|`nodus/evaluate`|
|Image|`nodus/evaluate:2.0.0`|
|Task|Evaluate|
|Modes|Evaluate|
|Launcher|Torchrun|
|Requires|model yes, data no, environment no|
|Checkpoints|`/nodus/state` (HFTrainer)|

## Parameters

Set under `spec.parameters`; the server validates them against this schema, and unset parameters take the runtime default.

|Parameter|Type|Default|Allowed|Description|
|-|-|-|-|-|
|`batchSize`|integer|`8`|1 to 1024|Examples per forward pass|
|`evaluationBatchSize`|integer|none|1 to 1024|Prompts generated together during Environment baseline and final evaluation|
|`limit`|integer|none|1 to 1000000|Samples per task (unset scores the whole task)|
|`maxCompletionLength`|integer|none|8 to 32768|Environment evaluation: tokens per completion|
|`numFewshot`|integer|none|0 to 64|Few-shot examples per prompt|
|`seed`|integer|`0`|0 to 2147483647|Seed for sampling and few-shot selection|
|`taskFilter`|object|none|at most 8 keys|Keep only tasks whose metadata has these values (such as family on reasoning-gym), before trainTasks and heldOutTasks count them|
|`tasks`|array of string|none|at most 64 items|Benchmark task names, such as arc_easy or hellaswag, whose datasets the run downloads from the Hugging Face Hub before it scores them offline; set either tasks or spec.environment, which scores an Environment’s held-out tasks instead|

## Outputs

|Output|Path|
|-|-|
|`outputs`|`/nodus/outputs`|

Every run also writes `results.json`, `provenance.json` and a sha256 `manifest.json`. Each file under `/nodus/outputs` is an output of the TrainingJob, named by its path: download one with `nodus cp tj/<name>:outputs/results.json ./results.json`.

Default resources: 1 × RTX-4090 or L40S or RTX-A6000 or A100-80G.

## Estimates

The dry-run estimate multiplies these by the run’s steps or tasks; the nightly runs re-measure them.

|Accelerator|Seconds|Per|
|-|-|-|
|A100-80G|0.17|secondsPerTask|
|H100|0.11|secondsPerTask|
|L40S|0.25|secondsPerTask|
|RTX-4090|0.3|secondsPerTask|
|RTX-A6000|0.33|secondsPerTask|

## Example

```yaml
apiVersion: nodus.dev/v1beta1
kind: TrainingJob
metadata: {name: my-evaluate}
spec:
  runtime: nodus/evaluate
  mode: Evaluate
  model: {uri: "hf://Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca"}
  parameters: {tasks: [arc_easy], limit: 200}
  maxCostUSD: "5.00"
```
