Training runtimes, Environments and the training SDK

A TrainingJob names a runtime, a model, data or an Environment, and the GPUs to use; the runtime does the rest. Eight catalog runtimes ship: sft (supervised fine-tuning, and pretraining from scratch with task: Pretrain), dpo, orpo, kto, reward-model, grpo-lora (RL with a baseline and a final evaluation on the same held-out tasks), distill (on-policy or offline distillation from a teacher) and evaluate. Each one validates its parameters before it loads a model; the training runtimes support LoRA and QLoRA and write the model or adapter with a sha256 manifest and its provenance. Set resources.nodes above 1 for a multi-node run with DDP, FSDP or ZeRO 1 to 3. --dry-run=server reviews the estimated time and cost before anything starts.

Five Environments grade model output for RL and evaluation: graph coloring, arithmetic, GSM8K, reasoning-gym and python-functions. Python-functions runs candidate code in a sandbox against private test cases. Each Environment keeps its training and held-out tasks apart, and its catalog examples record the measured baseline. Start one with nodus create trainingjob my-run --from-example nodus/gsm8k:gsm8k-trained.

In Python, nodus.recipes.finetune and nodus.recipes.rl build the same TrainingJobs (.preview() returns the estimate, .run() starts the job), and NodusCallback adds metrics, progress and checkpoints to your own Transformers or Lightning trainer. See the runtime reference and the Environment reference.