Skip to content

letter-counting-legacy-eval

View Markdown

Legacy letter-counting evaluation protocol: first 64 generated tasks, raw completion scoring.

Field Value
Reference nodus/letter-counting-legacy-eval@1.0.0
Image nodus/env-letter-counting-legacy-eval:1.0.0
Publisher Open-Thought
Category Reasoning
Readiness Research
Modes Evaluate
Reward Scalar
Held-out measures TrainedTask
Splits 0 train, 64 test (disjoint by canonical identity)
Licenses code Apache-2.0, data Apache-2.0
Source https://github.com/open-thought/reasoning-gym

Completions are graded by the platform, never by the trainer: the trainer submits {taskId, completion} batches and the verdicts come back as task events.

Grader Kind What it checks
score-answer Program Raw completion scored by pinned Reasoning Gym; only full credit is Correct

One line of nodus-env tasks --split test --seed 0; tasks never carry answers.

{
"metadata": {
"family": "letter_counting",
"systemPrompt": "Reply with the final answer only. Do not restate the question, do not show working, and do not add any words, labels or punctuation around it."
},
"prompt": "How many times does the letter \"b\" appear in the text: \"phrase Project Gutenberg associated with or appearing on the work you\"?",
"taskId": "letter-counting-legacy-eval:test:0:0"
}

Each example is a TrainingJob template. A baseline is shown only where the example was measured by running it.

Example Mode Runtime Model Tasks Baseline Trained Measured on
letter-counting Evaluate nodus/evaluate Qwen/Qwen3-1.7B @ 70d244c 64 not measured n/a (evaluation) not measured

Run one with a server dry-run first:

Terminal window
$ nodus create trainingjob my-run --from-example nodus/letter-counting-legacy-eval:letter-counting --dry-run=server -o estimate
import nodus
job = nodus.recipes.TrainingJob.from_example("nodus/letter-counting-legacy-eval:letter-counting")
plan = job.preview()
run = plan.run(max_cost=5)
print(run.wait().summary)