arithmetic-v2
View MarkdownEvaluate short integer expressions with standard precedence and give the exact result.
| Field | Value |
|---|---|
| Reference | nodus/arithmetic-v2@2.0.0 |
| Image | nodus/env-arithmetic-v2:2.0.0 |
| Publisher | Nodus |
| Category | Math |
| Readiness | Stable |
| Modes | Train, Evaluate |
| Reward | Binary |
| Held-out measures | TrainedTask |
| Splits | 500 train, 128 test (disjoint by canonical identity) |
| Licenses | code Apache-2.0, data Apache-2.0 |
| Source | https://github.com/nodus-compute/nodus-platform/tree/main/images/environments |
Grading
Section titled “Grading”Completions are graded by the platform, never by the trainer: the trainer submits {taskId, completion} batches and the verdicts come back as task events.
| Grader | Kind | What it checks |
|---|---|---|
exact-integer |
Program | The last |
Sample task
Section titled “Sample task”One line of nodus-env tasks --split test --seed 0; tasks never carry answers.
{ "metadata": { "terms": 3 }, "prompt": "Compute 41 - 38 * 33. Give only the final integer, like \u003canswer\u003e42\u003c/answer\u003e.", "taskId": "arithmetic-v2:test:0:0"}Examples
Section titled “Examples”Each example is a TrainingJob template. A baseline is shown only where the example was measured by running it.
| Example | Mode | Runtime | Model | Tasks | Baseline | Trained | Measured on |
|---|---|---|---|---|---|---|---|
arithmetic-grpo |
Train | nodus/grpo-lora |
Qwen/Qwen3-0.6B @ c1899de |
64 | not measured | not measured | not measured |
Run one with a server dry-run first:
$ nodus create trainingjob my-run --from-example nodus/arithmetic-v2:arithmetic-grpo --dry-run=server -o estimateimport nodus
job = nodus.recipes.TrainingJob.from_example("nodus/arithmetic-v2:arithmetic-grpo")plan = job.preview()run = plan.run(max_cost=5)print(run.wait().summary)