TrainingJob (nodus.dev/v1beta1, Beta) runs a training run on a TrainingRuntime and compiles it into a Job:
dry runs show the preset resources and an estimate from measured step times; parameters are checked against
the runtime’s schema; Hugging Face models are cached once per revision as read-only Volumes.
Managed runtimes for supervised fine-tuning (full, LoRA and QLoRA), continued and from-scratch pretraining,
DPO, ORPO, KTO, reward models, distillation, single-node GRPO and evaluation.
Multi-node training: resources.nodes above 1 runs a gang under torchrun (Accelerate and DeepSpeed configs
included) with distributed checkpoints and whole-gang restarts.
Environments (nodus.dev/v1, GA): graph coloring, arithmetic, GSM8K, reasoning-gym and python-functions,
graded by Nodus in isolated Sandboxes, with before-and-after comparisons that say when a change is measurable.
maxCostUSD on a TrainingJob caps its Job and its grading; raising it resumes a suspended run.