Skip to content

grpo-lora

View Markdown

GRPO reinforcement learning with a LoRA adapter over an Environment, graded by the platform.

Field Value
Reference nodus/grpo-lora
Image nodus/grpo-lora:2.0.0
Task GRPO
Modes Train, Evaluate
Launcher Torchrun
Requires model yes, data no, environment yes
Checkpoints /nodus/state (HFTrainer)

Set under spec.parameters; the server validates them against this schema, and unset parameters take the runtime default.

Parameter Type Default Allowed Description
attention string none eager, sdpa, flash_attention_2 Attention kernel
batchSize integer 1 1 to 1024 Examples per GPU per step
beta number 0 0 to 10 KL penalty toward the base model; 0 disables the reference model
distributed.findUnusedParameters boolean none DDP: tolerate parameters without gradients
distributed.strategy string Auto Auto, DDP, FSDP, ZeRO1, ZeRO2, ZeRO3 Auto and DDP replicate the model; FSDP and ZeRO shard it
epochs number 1 0.01 to 100 Passes over the data when steps is -1
evalSteps integer none 1 to 100000 Steps between validation passes; unset evaluates at the end only
evaluationBatchSize integer none 1 to 1024 Prompts generated together during Environment baseline and final evaluation
generationBatchSize integer none 2 to 65536 Total sampled completions per generation batch; must divide evenly across GPUs and numGenerations
gradientAccumulation integer 4 1 to 1024 Steps whose gradients are summed before an update
gradientCheckpointing boolean true Recompute activations to save GPU memory
learningRate number 0.000001 0 to 1 Peak learning rate
loggingSteps integer 10 1 to 10000 Steps between metric reports
lora.alpha integer 16 1 to 4096 Adapter scaling numerator
lora.dropout number 0.05 0 to 0.9 Adapter dropout
lora.r integer 8 1 to 1024 Adapter rank
lora.targetModules any none Module names, or all-linear
lrScheduler string cosine cosine, linear, constant, constant_with_warmup Learning-rate schedule
maxCompletionLength integer 64 8 to 32768 Tokens per completion; in a multi-turn episode, tokens per reply
maxEpisodeTokens integer none 8 to 131072 Tokens of a whole multi-turn episode after the prompt, the model’s and the Environment’s
maxGradNorm number none 0 to 100 Gradient clipping norm
maxLength integer 2048 16 to 131072 Tokens per example after truncation (packed block size for pretraining)
maxTurns integer none 1 to 31 Model turns per episode of a multi-turn Environment; 1 grades the first reply alone
method string LoRA Full, LoRA, QLoRA Full fine-tuning, a LoRA adapter, or LoRA on a 4-bit quantized base (GPU only)
numGenerations integer 4 2 to 64 Completions sampled per prompt (the GRPO group)
precision string Auto Auto, bf16, fp16, fp32 Numeric precision; Auto is bf16 where the GPU supports it, else fp16
sampling.temperature number 0.7 0 to 5 Sampling temperature
sampling.topK integer none 0 to 1000000 Keep the highest-probability K tokens; 0 disables top-k filtering
sampling.topP number 1 0 to 1 Nucleus sampling mass
saveSteps integer 25 1 to 100000 Steps between recovery checkpoints
saveTotalLimit integer none 1 to 10 Recovery checkpoints kept in the state directory
seed integer 42 0 to 2147483647 Seed for data order, splits and initialization
steps integer 50 -1 to 1000000 Optimizer steps; -1 trains for epochs instead
taskFilter object none at most 8 keys Keep only tasks whose metadata has these values (such as family on reasoning-gym), before trainTasks and heldOutTasks count them
warmupRatio number 0.03 0 to 0.5 Fraction of steps spent warming up the learning rate
weightDecay number none 0 to 1 AdamW weight decay
Output Path
outputs /nodus/outputs

Every run also writes results.json, provenance.json and a sha256 manifest.json. Each file under /nodus/outputs is an output of the TrainingJob, named by its path: download one with nodus cp tj/<name>:outputs/results.json ./results.json.

Without spec.resources, the TrainingJob takes the resources of the preset matching its model, method and quantization, else the runtime default.

Model Method Quantization GPUs
Qwen/Qwen3-0.6B LoRA none 1 × RTX-4090 or L40S or RTX-A6000
Qwen/Qwen3-1.7B LoRA none 1 × RTX-4090 or L40S or RTX-A6000
Qwen/Qwen3-4B QLoRA nf4 1 × RTX-4090 or L40S or RTX-A6000
Qwen/Qwen3-4B LoRA none 1 × L40S or RTX-A6000 or A100-80G
Qwen/Qwen3-8B QLoRA nf4 1 × L40S or RTX-A6000 or A100-80G

Default resources: 1 × RTX-4090 or L40S or RTX-A6000 or A100-80G.

The dry-run estimate multiplies these by the run’s steps or tasks; the nightly runs re-measure them.

Accelerator Seconds Per
A100-80G 0.7 secondsPerStep
H100 0.45 secondsPerStep
L40S 1 secondsPerStep
RTX-4090 1.2 secondsPerStep
RTX-A6000 1.3 secondsPerStep
apiVersion: nodus.dev/v1beta1
kind: TrainingJob
metadata: {name: my-grpo-lora}
spec:
runtime: nodus/grpo-lora
model: {uri: "hf://Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca"}
environment: {name: nodus/graph-coloring@1.0.0, trainTasks: 200, heldOutTasks: 64, seed: 42}
maxCostUSD: "5.00"