Training runtimes, Environments and the training SDK
A TrainingJob names a runtime, a model, data or an Environment, and the GPUs to use; the runtime does the rest.
Eight catalog runtimes ship:
sft (supervised fine-tuning, and pretraining from scratch with task: Pretrain), dpo, orpo,
kto, reward-model, grpo-lora (RL with a baseline and a final evaluation on the same held-out tasks),
distill (on-policy or offline distillation from a teacher) and evaluate. Each one validates its
parameters before it loads a model; the training runtimes support LoRA and QLoRA and write the model or adapter
with a sha256 manifest and its provenance. Set resources.nodes above 1 for a multi-node run with DDP, FSDP or ZeRO 1 to 3.
--dry-run=server reviews the estimated time and cost before anything starts.
Five Environments grade model output for RL and evaluation: graph coloring, arithmetic, GSM8K, reasoning-gym and
python-functions. Python-functions runs candidate code in a sandbox against private test cases. Each Environment
keeps its training and held-out tasks apart, and its catalog examples record the measured baseline. Start one with
nodus create trainingjob my-run --from-example nodus/gsm8k:gsm8k-trained.
In Python, nodus.recipes.finetune and nodus.recipes.rl build the same TrainingJobs (.preview() returns the
estimate, .run() starts the job), and NodusCallback adds metrics, progress and checkpoints to your own
Transformers or Lightning trainer. See the runtime reference and the
Environment reference.