reasoning-gym
View MarkdownProcedural reasoning generators with deterministic algorithmic scorers; only reviewed families are served.
| Field | Value |
|---|---|
| Reference | nodus/reasoning-gym@1.0.0 |
| Image | nodus/env-reasoning-gym:1.0.0 |
| Publisher | Open-Thought |
| Category | Reasoning |
| Readiness | Research |
| Modes | Train, Evaluate |
| Reward | Scalar |
| Held-out measures | TrainedTask |
| Splits | 2000 train, 320 test (disjoint by canonical identity) |
| Licenses | code Apache-2.0, data Apache-2.0 |
| Source | https://github.com/open-thought/reasoning-gym |
Grading
Section titled “Grading”Completions are graded by the platform, never by the trainer: the trainer submits {taskId, completion} batches and the verdicts come back as task events.
| Grader | Kind | What it checks |
|---|---|---|
score-answer |
Program | The family’s score_answer; only a full score is Correct, the score is the reward |
Sample task
Section titled “Sample task”One line of nodus-env tasks --split test --seed 0; tasks never carry answers.
{ "metadata": { "family": "basic_arithmetic" }, "prompt": "Calculate -3 - 8 / ( -8 + 1 + 8 ).\n\nReply with the final answer only. Do not restate the question, do not show working, and do not add any words, labels or punctuation around it.", "taskId": "reasoning-gym:test:0:0"}Examples
Section titled “Examples”Each example is a TrainingJob template. A baseline is shown only where the example was measured by running it.
| Example | Mode | Runtime | Model | Tasks | Baseline | Trained | Measured on |
|---|---|---|---|---|---|---|---|
chain-sum |
Evaluate | nodus/evaluate |
Qwen/Qwen3-1.7B @ 70d244c |
64 | not measured | n/a (evaluation) | not measured |
number-format |
Evaluate | nodus/evaluate |
Qwen/Qwen3-1.7B @ 70d244c |
64 | not measured | n/a (evaluation) | not measured |
basic-arithmetic |
Evaluate | nodus/evaluate |
Qwen/Qwen3-1.7B @ 70d244c |
64 | not measured | n/a (evaluation) | not measured |
products |
Evaluate | nodus/evaluate |
Qwen/Qwen3-1.7B @ 70d244c |
64 | not measured | n/a (evaluation) | not measured |
letter-counting |
Evaluate | nodus/evaluate |
Qwen/Qwen3-1.7B @ 70d244c |
64 | not measured | n/a (evaluation) | not measured |
Run one with a server dry-run first:
$ nodus create trainingjob my-run --from-example nodus/reasoning-gym:chain-sum --dry-run=server -o estimateimport nodus
job = nodus.recipes.TrainingJob.from_example("nodus/reasoning-gym:chain-sum")plan = job.preview()run = plan.run(max_cost=5)print(run.wait().summary)