# distill

> Knowledge distillation from a teacher model; lmbda 0 is offline distillation

Source: https://nodus-platform-site.pages.dev/docs/reference/runtimes/distill/
Build revision: 211ad9f836655b1c3a2668c4693e442471f28614

<!-- Generated by tools/catalogdocs from catalog/objects. Do not edit: run make gen. -->

Beta

TrainingJob and TrainingRuntime are Beta: fields can still change before GA.

Knowledge distillation from a teacher model; lmbda 0 is offline distillation.

|Field|Value|
|-|-|
|Reference|`nodus/distill`|
|Image|`nodus/distill:2.0.0`|
|Task|Distill|
|Modes|Train|
|Launcher|Torchrun|
|Requires|model yes, data yes, environment no|
|Checkpoints|`/nodus/state` (HFTrainer)|

## Parameters

Set under `spec.parameters`; the server validates them against this schema, and unset parameters take the runtime default.

|Parameter|Type|Default|Allowed|Description|
|-|-|-|-|-|
|`alpha`|number|`0.5`|0 to 1|Generalized JSD interpolation: 0 forward KL, 1 reverse KL|
|`attention`|string|none|`eager`, `sdpa`, `flash_attention_2`|Attention kernel|
|`batchSize`|integer|`4`|1 to 1024|Examples per GPU per step|
|`distributed.findUnusedParameters`|boolean|none||DDP: tolerate parameters without gradients|
|`distributed.strategy`|string|`Auto`|`Auto`, `DDP`, `FSDP`, `ZeRO1`, `ZeRO2`, `ZeRO3`|Auto and DDP replicate the model; FSDP and ZeRO shard it|
|`epochs`|number|`1`|0.01 to 100|Passes over the data when `steps` is -1|
|`evalSteps`|integer|none|1 to 100000|Steps between validation passes; unset evaluates at the end only|
|`gradientAccumulation`|integer|`4`|1 to 1024|Steps whose gradients are summed before an update|
|`gradientCheckpointing`|boolean|`true`||Recompute activations to save GPU memory|
|`learningRate`|number|`0.00001`|0 to 1|Peak learning rate|
|`lmbda`|number|`0`|0 to 1|Fraction of batches whose completions the student samples itself; 0 is offline distillation|
|`loggingSteps`|integer|`10`|1 to 10000|Steps between metric reports|
|`lora.alpha`|integer|`32`|1 to 4096|Adapter scaling numerator|
|`lora.dropout`|number|`0.05`|0 to 0.9|Adapter dropout|
|`lora.r`|integer|`16`|1 to 1024|Adapter rank|
|`lora.targetModules`|any|none||Module names, or all-linear|
|`lrScheduler`|string|`cosine`|`cosine`, `linear`, `constant`, `constant_with_warmup`|Learning-rate schedule|
|`maxGradNorm`|number|none|0 to 100|Gradient clipping norm|
|`maxLength`|integer|`1024`|16 to 131072|Tokens per example after truncation (packed block size for pretraining)|
|`maxNewTokens`|integer|`128`|1 to 8192|Most new tokens per student-sampled completion|
|`method`|string|`LoRA`|`Full`, `LoRA`, `QLoRA`|Full fine-tuning, a LoRA adapter, or LoRA on a 4-bit quantized base (GPU only)|
|`precision`|string|`Auto`|`Auto`, `bf16`, `fp16`, `fp32`|Numeric precision; Auto is bf16 where the GPU supports it, else fp16|
|`saveSteps`|integer|`200`|1 to 100000|Steps between recovery checkpoints|
|`saveTotalLimit`|integer|none|1 to 10|Recovery checkpoints kept in the state directory|
|`seed`|integer|`42`|0 to 2147483647|Seed for data order, splits and initialization|
|`sequenceKD`|boolean|none||Train on teacher-generated sequences (sequence-level KD)|
|`steps`|integer|`-1`|-1 to 1000000|Optimizer steps; -1 trains for `epochs` instead|
|`temperature`|number|`2`|0.05 to 20|Softmax temperature of teacher and student|
|`warmupRatio`|number|`0.03`|0 to 0.5|Fraction of steps spent warming up the learning rate|
|`weightDecay`|number|none|0 to 1|AdamW weight decay|

## Outputs

|Output|Path|
|-|-|
|`outputs`|`/nodus/outputs`|

Every run also writes `results.json`, `provenance.json` and a sha256 `manifest.json`. Each file under `/nodus/outputs` is an output of the TrainingJob, named by its path: download one with `nodus cp tj/<name>:outputs/results.json ./results.json`.

## Presets

Without `spec.resources`, the TrainingJob takes the resources of the preset matching its model, method and quantization, else the runtime default.

|Model|Method|Quantization|GPUs|
|-|-|-|-|
|Qwen/Qwen3-0.6B|LoRA|none|1 × RTX-4090 or L40S or RTX-A6000|
|Qwen/Qwen3-0.6B|Full|none|1 × L40S or RTX-A6000 or A100-80G|
|Qwen/Qwen3-1.7B|LoRA|none|1 × RTX-4090 or L40S or RTX-A6000|
|Qwen/Qwen3-1.7B|Full|none|1 × A100-80G or H100|
|Qwen/Qwen3-4B|QLoRA|nf4|1 × RTX-4090 or L40S or RTX-A6000|
|Qwen/Qwen3-4B|LoRA|none|1 × L40S or RTX-A6000 or A100-80G|
|Qwen/Qwen3-8B|QLoRA|nf4|1 × L40S or RTX-A6000 or A100-80G|
|Qwen/Qwen3-8B|Full|none|8 × H100|

Default resources: 1 × RTX-4090 or L40S or RTX-A6000 or A100-80G.

## Estimates

The dry-run estimate multiplies these by the run’s steps or tasks; the nightly runs re-measure them.

|Accelerator|Seconds|Per|
|-|-|-|
|A100-80G|0.7|secondsPerStep|
|H100|0.45|secondsPerStep|
|L40S|1|secondsPerStep|
|RTX-4090|1.2|secondsPerStep|
|RTX-A6000|1.3|secondsPerStep|

## Example

```yaml
apiVersion: nodus.dev/v1beta1
kind: TrainingJob
metadata: {name: my-distill}
spec:
  runtime: nodus/distill
  model: {uri: "hf://Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca"}
  data: {volume: my-dataset, format: JSONL}
  maxCostUSD: "5.00"
```
