Train on a reward function in one call
rl.train(tasks, reward) is the shortest path from a task and a reward to a trained adapter, and
python -m nodus.examples.rl runs a whole example in one command. The base model is Qwen/Qwen3-0.6B unless you pass
model=. The
SDK packages your reward, and the lines of its file it uses, as an Environment in your project. It holds out a fifth
of the tasks for the before-and-after comparison and runs GRPO with LoRA. run.watch() prints the reward, loss and
KL of each step as it trains, and run.tasks() lists each graded task with its outcome and reward.
Each step learns from groups of replies: 8 replies to each of 4 tasks, sampled at temperature 1, with replies of up to
256 tokens and a 4e-5 learning rate. rl.train("nodus/gsm8k@1.0.0")
runs a catalog Environment with the same defaults.
The console now only shows training: runs, curves, gradings, results, cost and outputs. The New training, New environment and Start training screens are gone; runs start from the SDK or the CLI.