dpo
View MarkdownDirect preference optimization on prompt, chosen and rejected rows.
| Field | Value |
|---|---|
| Reference | nodus/dpo |
| Image | nodus/dpo:2.0.0 |
| Task | DPO |
| Modes | Train |
| Launcher | Torchrun |
| Requires | model yes, data yes, environment no |
| Checkpoints | /nodus/state (HFTrainer) |
Parameters
Section titled “Parameters”Set under spec.parameters; the server validates them against this schema, and unset parameters take the runtime default.
| Parameter | Type | Default | Allowed | Description |
|---|---|---|---|---|
attention |
string | none | eager, sdpa, flash_attention_2 |
Attention kernel |
batchSize |
integer | 4 |
1 to 1024 | Examples per GPU per step |
beta |
number | 0.1 |
0 to 10 | Strength of the KL penalty toward the reference model |
distributed.findUnusedParameters |
boolean | none | DDP: tolerate parameters without gradients | |
distributed.strategy |
string | Auto |
Auto, DDP, FSDP, ZeRO1, ZeRO2, ZeRO3 |
Auto and DDP replicate the model; FSDP and ZeRO shard it |
epochs |
number | 1 |
0.01 to 100 | Passes over the data when steps is -1 |
evalSteps |
integer | none | 1 to 100000 | Steps between validation passes; unset evaluates at the end only |
gradientAccumulation |
integer | 4 |
1 to 1024 | Steps whose gradients are summed before an update |
gradientCheckpointing |
boolean | true |
Recompute activations to save GPU memory | |
learningRate |
number | 0.000005 |
0 to 1 | Peak learning rate |
loggingSteps |
integer | 10 |
1 to 10000 | Steps between metric reports |
lora.alpha |
integer | 32 |
1 to 4096 | Adapter scaling numerator |
lora.dropout |
number | 0.05 |
0 to 0.9 | Adapter dropout |
lora.r |
integer | 16 |
1 to 1024 | Adapter rank |
lora.targetModules |
any | none | Module names, or all-linear | |
lossType |
string | sigmoid |
sigmoid, ipo, hinge, robust, sppo_hard, nca_pair, apo_zero, apo_down |
DPO loss |
lrScheduler |
string | cosine |
cosine, linear, constant, constant_with_warmup |
Learning-rate schedule |
maxGradNorm |
number | none | 0 to 100 | Gradient clipping norm |
maxLength |
integer | 1024 |
16 to 131072 | Tokens per example after truncation (packed block size for pretraining) |
method |
string | LoRA |
Full, LoRA, QLoRA |
Full fine-tuning, a LoRA adapter, or LoRA on a 4-bit quantized base (GPU only) |
precision |
string | Auto |
Auto, bf16, fp16, fp32 |
Numeric precision; Auto is bf16 where the GPU supports it, else fp16 |
saveSteps |
integer | 200 |
1 to 100000 | Steps between recovery checkpoints |
saveTotalLimit |
integer | none | 1 to 10 | Recovery checkpoints kept in the state directory |
seed |
integer | 42 |
0 to 2147483647 | Seed for data order, splits and initialization |
steps |
integer | -1 |
-1 to 1000000 | Optimizer steps; -1 trains for epochs instead |
warmupRatio |
number | 0.03 |
0 to 0.5 | Fraction of steps spent warming up the learning rate |
weightDecay |
number | none | 0 to 1 | AdamW weight decay |
Outputs
Section titled “Outputs”| Output | Path |
|---|---|
outputs |
/nodus/outputs |
Every run also writes results.json, provenance.json and a sha256 manifest.json. Each file under /nodus/outputs is an output of the TrainingJob, named by its path: download one with nodus cp tj/<name>:outputs/results.json ./results.json.
Presets
Section titled “Presets”Without spec.resources, the TrainingJob takes the resources of the preset matching its model, method and quantization, else the runtime default.
| Model | Method | Quantization | GPUs |
|---|---|---|---|
| Qwen/Qwen3-0.6B | LoRA | none | 1 × RTX-4090 or L40S or RTX-A6000 |
| Qwen/Qwen3-0.6B | Full | none | 1 × L40S or RTX-A6000 or A100-80G |
| Qwen/Qwen3-1.7B | LoRA | none | 1 × RTX-4090 or L40S or RTX-A6000 |
| Qwen/Qwen3-1.7B | Full | none | 1 × A100-80G or H100 |
| Qwen/Qwen3-4B | QLoRA | nf4 | 1 × RTX-4090 or L40S or RTX-A6000 |
| Qwen/Qwen3-4B | LoRA | none | 1 × L40S or RTX-A6000 or A100-80G |
| Qwen/Qwen3-8B | QLoRA | nf4 | 1 × L40S or RTX-A6000 or A100-80G |
| Qwen/Qwen3-8B | Full | none | 8 × H100 |
Default resources: 1 × RTX-4090 or L40S or RTX-A6000 or A100-80G.
Estimates
Section titled “Estimates”The dry-run estimate multiplies these by the run’s steps or tasks; the nightly runs re-measure them.
| Accelerator | Seconds | Per |
|---|---|---|
| A100-80G | 0.7 | secondsPerStep |
| H100 | 0.45 | secondsPerStep |
| L40S | 1 | secondsPerStep |
| RTX-4090 | 1.2 | secondsPerStep |
| RTX-A6000 | 1.3 | secondsPerStep |
Example
Section titled “Example”apiVersion: nodus.dev/v1beta1kind: TrainingJobmetadata: {name: my-dpo}spec: runtime: nodus/dpo model: {uri: "hf://Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca"} data: {volume: my-dataset, format: JSONL} maxCostUSD: "5.00"