# nodus.recipes.rl

> Reinforcement learning and evaluation on catalog Environments or your own.

Source: https://nodus-platform-site.pages.dev/docs/reference/python/nodus-recipes-rl/
Build revision: 211ad9f836655b1c3a2668c4693e442471f28614

<!-- Generated by scripts/gen-reference.mjs from the SDK docstrings (griffe). Do not edit: run make gen. -->

Reinforcement learning and evaluation on catalog Environments or your own.

From a task list and a reward function to a training run in one call. `python -m nodus.examples.rl` runs a whole example (`nodus/examples/rl.py`), a template to copy:

```plaintext
def reward(completion, answer):
    return 1.0 if answer in completion else 0.0

run = rl.train([("Spell 'cat' backwards.", "tac"), ...], reward, max_cost=2)   # model= picks another base model
run.watch()                                           # each stage, then reward, loss and KL per step
run.outputs.download("./outputs")                     # the LoRA adapter and the before/after comparison
```

`train` packages the reward with the statements of its file it uses as an Environment in your project, then runs GRPO on it, learning from groups of replies. `rl.train("nodus/gsm8k@1.0.0")` trains on a catalog Environment the same way. To set every GRPO parameter yourself, on a catalog Environment or one you build (`examples/training/custom-reward`):

```plaintext
job = rl.grpo_lora(model="Qwen/Qwen3-0.6B", environment="nodus/graph-coloring@1.0.0",
                   steps=50, held_out=64, seed=42)
plan = job.preview()                                  # estimate, compiled Job, blocking reasons, ETag
run = plan.run(gpu="RTX-4090", max_cost=5, idempotency_key="gc-001")
print(run.wait().summary.comparison)                  # the numbers the console shows
```

The trainer never grades itself: completions go to the platform grader and come back as task events, and the comparison is computed server-side on the same held-out tasks before and after training.

## `evaluate`

```python
evaluate(*, model: Any, environment: str | None = None, tasks: Iterable[str] | None = None, held_out: int = 64, seed: int = 0, task_filter: Mapping[str, str] | None = None, **params: Any) -> Any
```

`mode: Evaluate` on `evaluate`: an Environment’s held-out tasks, or benchmark `tasks` such as arc_easy.

## `grpo_lora`

```python
grpo_lora(*, model: Any, environment: str, steps: int = 50, train_tasks: int = 256, held_out: int = 64, seed: int = 0, lora: LoRA | None = None, task_filter: Mapping[str, str] | None = None, min_comparison_tasks: int = 16, **params: Any) -> Any
```

GRPO with a LoRA adapter (`grpo-lora`): baseline, train, then the final eval on the same held-out tasks.
