Training runtimes
View MarkdownThe catalog runtimes in project nodus; reference one as spec.runtime: nodus/<name> on a TrainingJob.
| Runtime | Task | Modes | Summary |
|---|---|---|---|
distill |
Distill | Train | Knowledge distillation from a teacher model; lmbda 0 is offline distillation |
dpo |
DPO | Train | Direct preference optimization on prompt, chosen and rejected rows |
evaluate |
Evaluate | Evaluate | Benchmark evaluation on standard tasks, or over an Environment’s held-out tasks |
grpo-lora |
GRPO | Train, Evaluate | GRPO reinforcement learning with a LoRA adapter over an Environment, graded by the platform |
kto |
KTO | Train | Kahneman-Tversky optimization on unpaired desirable and undesirable completions |
orpo |
ORPO | Train | Odds-ratio preference optimization: SFT and preference alignment in one pass, no reference model |
reward-model |
Reward | Train | Reward-model training: a one-logit head scored on chosen versus rejected |
sft |
SFT | Train | Supervised fine-tuning (full, LoRA, QLoRA) and pretraining with task: Pretrain |