Train on a reward function in one call

rl.train(tasks, reward) is the shortest path from a task and a reward to a trained adapter, and python -m nodus.examples.rl runs a whole example in one command. The base model is Qwen/Qwen3-0.6B unless you pass model=. The SDK packages your reward, and the lines of its file it uses, as an Environment in your project. It holds out a fifth of the tasks for the before-and-after comparison and runs GRPO with LoRA. run.watch() prints the reward, loss and KL of each step as it trains, and run.tasks() lists each graded task with its outcome and reward.

Each step learns from groups of replies: 8 replies to each of 4 tasks, sampled at temperature 1, with replies of up to 256 tokens and a 4e-5 learning rate. rl.train("nodus/gsm8k@1.0.0") runs a catalog Environment with the same defaults.

The console now only shows training: runs, curves, gradings, results, cost and outputs. The New training, New environment and Start training screens are gone; runs start from the SDK or the CLI.