Skip to content

Eight thousand grade-school maths word problems, each with a worked solution and a single numeric answer.

Field Value
Reference nodus/gsm8k@1.0.0
Image nodus/env-gsm8k:1.0.0
Publisher OpenAI
Category Math
Readiness Research
Modes Train, Evaluate
Reward Binary
Held-out measures TrainedTask
Splits 7473 train, 1319 test (disjoint by canonical identity)
Licenses code MIT, data MIT
Source https://huggingface.co/datasets/openai/gsm8k/tree/740312add88f781978c0658806c59bc2815b9866

Completions are graded by the platform, never by the trainer: the trainer submits {taskId, completion} batches and the verdicts come back as task events.

Grader Kind What it checks
final-number Program The last number, else the last number, equals the value after ####

One line of nodus-env tasks --split test --seed 0; tasks never carry answers.

{
"metadata": {
"source": "openai/gsm8k"
},
"prompt": "A fruit vendor bought 50 watermelons for $80. He sold all of them at a profit of 25%. How much was each watermelon sold?\n\nWork through the problem step by step, then give the final number inside answer tags, like \u003canswer\u003e42\u003c/answer\u003e.",
"taskId": "gsm8k:test:0:0"
}

Each example is a TrainingJob template. A baseline is shown only where the example was measured by running it.

Example Mode Runtime Model Tasks Baseline Trained Measured on
gsm8k-trained Train nodus/grpo-lora Qwen/Qwen3-1.7B @ 70d244c 64 79.7 % 87.5 % a40-48g-x1, 2026-09-24

Run one with a server dry-run first:

Terminal window
$ nodus create trainingjob my-run --from-example nodus/gsm8k:gsm8k-trained --dry-run=server -o estimate
import nodus
job = nodus.recipes.TrainingJob.from_example("nodus/gsm8k:gsm8k-trained")
plan = job.preview()
run = plan.run(max_cost=5)
print(run.wait().summary)