nodus.recipes.rl
View MarkdownReinforcement learning and evaluation on catalog Environments or your own.
From a task list and a reward function to a training run in one call. python -m nodus.examples.rl runs a whole
example (nodus/examples/rl.py), a template to copy:
def reward(completion, answer): return 1.0 if answer in completion else 0.0
run = rl.train([("Spell 'cat' backwards.", "tac"), ...], reward, max_cost=2) # model= picks another base modelrun.watch() # each stage, then reward, loss and KL per steprun.outputs.download("./outputs") # the LoRA adapter and the before/after comparisontrain packages the reward with the statements of its file it uses as an Environment in your project, then runs
GRPO on it, learning from groups of replies. rl.train("nodus/gsm8k@1.0.0") trains on a catalog Environment the
same way. To set every GRPO parameter yourself, on a catalog Environment or one you build
(examples/training/custom-reward):
job = rl.grpo_lora(model="Qwen/Qwen3-0.6B", environment="nodus/graph-coloring@1.0.0", steps=50, held_out=64, seed=42)plan = job.preview() # estimate, compiled Job, blocking reasons, ETagrun = plan.run(gpu="RTX-4090", max_cost=5, idempotency_key="gc-001")print(run.wait().summary.comparison) # the numbers the console showsThe trainer never grades itself: completions go to the platform grader and come back as task events, and the comparison is computed server-side on the same held-out tasks before and after training.
evaluate
Section titled “evaluate”evaluate(*, model: Any, environment: str | None = None, tasks: Iterable[str] | None = None, held_out: int = 64, seed: int = 0, task_filter: Mapping[str, str] | None = None, **params: Any) -> Anymode: Evaluate on evaluate: an Environment’s held-out tasks, or benchmark tasks such as arc_easy.
grpo_lora
Section titled “grpo_lora”grpo_lora(*, model: Any, environment: str, steps: int = 50, train_tasks: int = 256, held_out: int = 64, seed: int = 0, lora: LoRA | None = None, task_filter: Mapping[str, str] | None = None, min_comparison_tasks: int = 16, **params: Any) -> AnyGRPO with a LoRA adapter (grpo-lora): baseline, train, then the final eval on the same held-out tasks.