# nodus.recipes.finetune

> Fine-tuning, preference training, distillation and pretraining recipes on the catalog runtimes.

Source: https://nodus-platform-site.pages.dev/docs/reference/python/nodus-recipes-finetune/
Build revision: 211ad9f836655b1c3a2668c4693e442471f28614

<!-- Generated by scripts/gen-reference.mjs from the SDK docstrings (griffe). Do not edit: run make gen. -->

Fine-tuning, preference training, distillation and pretraining recipes on the catalog runtimes.

Each function returns a `TrainingJob` builder; nothing runs until `.preview()` or `.run(gpu=…, max_cost=…)`. Keyword arguments not named here are runtime parameters in snake_case (`learning_rate=1e-5` is `parameters.learningRate`), validated by the server against the runtime’s schema.

```plaintext
sft = finetune.sft(model="Qwen/Qwen3-1.7B", data=nodus.Volume.from_name("support-chats"),
                   lora=finetune.LoRA(r=16), max_steps=2000).run(gpu="H100", max_cost=20)
pt = finetune.pretrain(model="Qwen/Qwen3-0.6B", data=corpus, initialization="Scratch")
big = finetune.sft(model=m, data=d, distributed="ZeRO3").run(gpu="H100:8", nodes=2, max_cost=200)
```

## Exports

* `Data`: defined in `nodus.recipes._job`
* `LoRA`: defined in `nodus.recipes._job`

## `distill`

```python
distill(*, student: Any, teacher: Any, data: Any, temperature: float | None = None, alpha: float | None = None, lmbda: float | None = None, **params: Any) -> Any
```

Knowledge distillation from a teacher model (`distill`); `lmbda=0` is offline KD on the dataset’s text.

## `dpo`

```python
dpo(*, model: Any, data: Any, beta: float | None = None, **params: Any) -> Any
```

Direct preference optimisation on prompt, chosen and rejected rows (`dpo`).

## `kto`

```python
kto(*, model: Any, data: Any, beta: float | None = None, **params: Any) -> Any
```

KTO on prompt, completion and a boolean label (`kto`).

## `orpo`

```python
orpo(*, model: Any, data: Any, beta: float | None = None, **params: Any) -> Any
```

Odds-ratio preference optimisation, no reference model (`orpo`).

## `pretrain`

```python
pretrain(*, model: Any, data: Any, initialization: str = 'Continued', packing: bool = True, **params: Any) -> Any
```

Causal-LM pretraining on raw text (`sft`, `task: Pretrain`).

`initialization="Continued"` keeps the model’s weights; `"Scratch"` uses only its config and tokenizer (optionally resized by `architecture={...}`) and starts from random weights.

## `reward`

```python
reward(*, model: Any, data: Any, **params: Any) -> Any
```

A reward model from chosen and rejected pairs (`reward-model`).

## `sft`

```python
sft(*, model: Any, data: Any, lora: LoRA | None = None, max_steps: int | None = None, **params: Any) -> Any
```

Supervised fine-tuning on prompt and completion (or chat `messages`) rows (`sft`).
