Skip to content

nodus.recipes.rl

View Markdown

Reinforcement learning and evaluation on catalog Environments or your own.

From a task list and a reward function to a training run in one call. python -m nodus.examples.rl runs a whole example (nodus/examples/rl.py), a template to copy:

def reward(completion, answer):
return 1.0 if answer in completion else 0.0
run = rl.train([("Spell 'cat' backwards.", "tac"), ...], reward, max_cost=2) # model= picks another base model
run.watch() # each stage, then reward, loss and KL per step
run.outputs.download("./outputs") # the LoRA adapter and the before/after comparison

train packages the reward with the statements of its file it uses as an Environment in your project, then runs GRPO on it, learning from groups of replies. rl.train("nodus/gsm8k@1.0.0") trains on a catalog Environment the same way. To set every GRPO parameter yourself, on a catalog Environment or one you build (examples/training/custom-reward):

job = rl.grpo_lora(model="Qwen/Qwen3-0.6B", environment="nodus/graph-coloring@1.0.0",
steps=50, held_out=64, seed=42)
plan = job.preview() # estimate, compiled Job, blocking reasons, ETag
run = plan.run(gpu="RTX-4090", max_cost=5, idempotency_key="gc-001")
print(run.wait().summary.comparison) # the numbers the console shows

The trainer never grades itself: completions go to the platform grader and come back as task events, and the comparison is computed server-side on the same held-out tasks before and after training.

evaluate(*, model: Any, environment: str | None = None, tasks: Iterable[str] | None = None, held_out: int = 64, seed: int = 0, task_filter: Mapping[str, str] | None = None, **params: Any) -> Any

mode: Evaluate on evaluate: an Environment’s held-out tasks, or benchmark tasks such as arc_easy.

grpo_lora(*, model: Any, environment: str, steps: int = 50, train_tasks: int = 256, held_out: int = 64, seed: int = 0, lora: LoRA | None = None, task_filter: Mapping[str, str] | None = None, min_comparison_tasks: int = 16, **params: Any) -> Any

GRPO with a LoRA adapter (grpo-lora): baseline, train, then the final eval on the same held-out tasks.