Skip to content

arithmetic-v2

View Markdown

Evaluate short integer expressions with standard precedence and give the exact result.

Field Value
Reference nodus/arithmetic-v2@2.0.0
Image nodus/env-arithmetic-v2:2.0.0
Publisher Nodus
Category Math
Readiness Stable
Modes Train, Evaluate
Reward Binary
Held-out measures TrainedTask
Splits 500 train, 128 test (disjoint by canonical identity)
Licenses code Apache-2.0, data Apache-2.0
Source https://github.com/nodus-compute/nodus-platform/tree/main/images/environments

Completions are graded by the platform, never by the trainer: the trainer submits {taskId, completion} batches and the verdicts come back as task events.

Grader Kind What it checks
exact-integer Program The last N equals the expression’s value

One line of nodus-env tasks --split test --seed 0; tasks never carry answers.

{
"metadata": {
"terms": 3
},
"prompt": "Compute 41 - 38 * 33. Give only the final integer, like \u003canswer\u003e42\u003c/answer\u003e.",
"taskId": "arithmetic-v2:test:0:0"
}

Each example is a TrainingJob template. A baseline is shown only where the example was measured by running it.

Example Mode Runtime Model Tasks Baseline Trained Measured on
arithmetic-grpo Train nodus/grpo-lora Qwen/Qwen3-0.6B @ c1899de 64 not measured not measured not measured

Run one with a server dry-run first:

Terminal window
$ nodus create trainingjob my-run --from-example nodus/arithmetic-v2:arithmetic-grpo --dry-run=server -o estimate
import nodus
job = nodus.recipes.TrainingJob.from_example("nodus/arithmetic-v2:arithmetic-grpo")
plan = job.preview()
run = plan.run(max_cost=5)
print(run.wait().summary)