# Training runtimes

> Every catalog TrainingRuntime, its task and modes.

Source: https://nodus-platform-site.pages.dev/docs/reference/runtimes/
Build revision: 211ad9f836655b1c3a2668c4693e442471f28614

<!-- Generated by tools/catalogdocs from catalog/objects. Do not edit: run make gen. -->

Beta

TrainingRuntimes are Beta.

The catalog runtimes in project `nodus`; reference one as `spec.runtime: nodus/<name>` on a TrainingJob.

|Runtime|Task|Modes|Summary|
|-|-|-|-|
|[`distill`](https://nodus-platform-site.pages.dev/docs/reference/runtimes/distill/)|Distill|Train|Knowledge distillation from a teacher model; lmbda 0 is offline distillation|
|[`dpo`](https://nodus-platform-site.pages.dev/docs/reference/runtimes/dpo/)|DPO|Train|Direct preference optimization on prompt, chosen and rejected rows|
|[`evaluate`](https://nodus-platform-site.pages.dev/docs/reference/runtimes/evaluate/)|Evaluate|Evaluate|Benchmark evaluation on standard tasks, or over an Environment’s held-out tasks|
|[`grpo-lora`](https://nodus-platform-site.pages.dev/docs/reference/runtimes/grpo-lora/)|GRPO|Train, Evaluate|GRPO reinforcement learning with a LoRA adapter over an Environment, graded by the platform|
|[`kto`](https://nodus-platform-site.pages.dev/docs/reference/runtimes/kto/)|KTO|Train|Kahneman-Tversky optimization on unpaired desirable and undesirable completions|
|[`orpo`](https://nodus-platform-site.pages.dev/docs/reference/runtimes/orpo/)|ORPO|Train|Odds-ratio preference optimization: SFT and preference alignment in one pass, no reference model|
|[`reward-model`](https://nodus-platform-site.pages.dev/docs/reference/runtimes/reward-model/)|Reward|Train|Reward-model training: a one-logit head scored on chosen versus rejected|
|[`sft`](https://nodus-platform-site.pages.dev/docs/reference/runtimes/sft/)|SFT|Train|Supervised fine-tuning (full, LoRA, QLoRA) and pretraining with task: Pretrain|
