Skip to content

reward-model

View Markdown

Reward-model training: a one-logit head scored on chosen versus rejected.

Field Value
Reference nodus/reward-model
Image nodus/reward-model:2.0.0
Task Reward
Modes Train
Launcher Torchrun
Requires model yes, data yes, environment no
Checkpoints /nodus/state (HFTrainer)

Set under spec.parameters; the server validates them against this schema, and unset parameters take the runtime default.

Parameter Type Default Allowed Description
attention string none eager, sdpa, flash_attention_2 Attention kernel
batchSize integer 4 1 to 1024 Examples per GPU per step
distributed.findUnusedParameters boolean none DDP: tolerate parameters without gradients
distributed.strategy string Auto Auto, DDP, FSDP, ZeRO1, ZeRO2, ZeRO3 Auto and DDP replicate the model; FSDP and ZeRO shard it
epochs number 1 0.01 to 100 Passes over the data when steps is -1
evalSteps integer none 1 to 100000 Steps between validation passes; unset evaluates at the end only
gradientAccumulation integer 4 1 to 1024 Steps whose gradients are summed before an update
gradientCheckpointing boolean true Recompute activations to save GPU memory
learningRate number 0.00001 0 to 1 Peak learning rate
loggingSteps integer 10 1 to 10000 Steps between metric reports
lora.alpha integer 32 1 to 4096 Adapter scaling numerator
lora.dropout number 0.05 0 to 0.9 Adapter dropout
lora.r integer 16 1 to 1024 Adapter rank
lora.targetModules any none Module names, or all-linear
lrScheduler string cosine cosine, linear, constant, constant_with_warmup Learning-rate schedule
maxGradNorm number none 0 to 100 Gradient clipping norm
maxLength integer 1024 16 to 131072 Tokens per example after truncation (packed block size for pretraining)
method string LoRA Full, LoRA, QLoRA Full fine-tuning, a LoRA adapter, or LoRA on a 4-bit quantized base (GPU only)
precision string Auto Auto, bf16, fp16, fp32 Numeric precision; Auto is bf16 where the GPU supports it, else fp16
saveSteps integer 200 1 to 100000 Steps between recovery checkpoints
saveTotalLimit integer none 1 to 10 Recovery checkpoints kept in the state directory
seed integer 42 0 to 2147483647 Seed for data order, splits and initialization
steps integer -1 -1 to 1000000 Optimizer steps; -1 trains for epochs instead
warmupRatio number 0.03 0 to 0.5 Fraction of steps spent warming up the learning rate
weightDecay number none 0 to 1 AdamW weight decay
Output Path
outputs /nodus/outputs

Every run also writes results.json, provenance.json and a sha256 manifest.json. Each file under /nodus/outputs is an output of the TrainingJob, named by its path: download one with nodus cp tj/<name>:outputs/results.json ./results.json.

Without spec.resources, the TrainingJob takes the resources of the preset matching its model, method and quantization, else the runtime default.

Model Method Quantization GPUs
Qwen/Qwen3-0.6B LoRA none 1 × RTX-4090 or L40S or RTX-A6000
Qwen/Qwen3-0.6B Full none 1 × L40S or RTX-A6000 or A100-80G
Qwen/Qwen3-1.7B LoRA none 1 × RTX-4090 or L40S or RTX-A6000
Qwen/Qwen3-1.7B Full none 1 × A100-80G or H100
Qwen/Qwen3-4B QLoRA nf4 1 × RTX-4090 or L40S or RTX-A6000
Qwen/Qwen3-4B LoRA none 1 × L40S or RTX-A6000 or A100-80G
Qwen/Qwen3-8B QLoRA nf4 1 × L40S or RTX-A6000 or A100-80G
Qwen/Qwen3-8B Full none 8 × H100

Default resources: 1 × RTX-4090 or L40S or RTX-A6000 or A100-80G.

The dry-run estimate multiplies these by the run’s steps or tasks; the nightly runs re-measure them.

Accelerator Seconds Per
A100-80G 0.7 secondsPerStep
H100 0.45 secondsPerStep
L40S 1 secondsPerStep
RTX-4090 1.2 secondsPerStep
RTX-A6000 1.3 secondsPerStep
apiVersion: nodus.dev/v1beta1
kind: TrainingJob
metadata: {name: my-reward-model}
spec:
runtime: nodus/reward-model
model: {uri: "hf://Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca"}
data: {volume: my-dataset, format: JSONL}
maxCostUSD: "5.00"