grpo-lora
View MarkdownGRPO reinforcement learning with a LoRA adapter over an Environment, graded by the platform.
| Field | Value |
|---|---|
| Reference | nodus/grpo-lora |
| Image | nodus/grpo-lora:2.0.0 |
| Task | GRPO |
| Modes | Train, Evaluate |
| Launcher | Torchrun |
| Requires | model yes, data no, environment yes |
| Checkpoints | /nodus/state (HFTrainer) |
Parameters
Section titled “Parameters”Set under spec.parameters; the server validates them against this schema, and unset parameters take the runtime default.
| Parameter | Type | Default | Allowed | Description |
|---|---|---|---|---|
attention |
string | none | eager, sdpa, flash_attention_2 |
Attention kernel |
batchSize |
integer | 1 |
1 to 1024 | Examples per GPU per step |
beta |
number | 0 |
0 to 10 | KL penalty toward the base model; 0 disables the reference model |
distributed.findUnusedParameters |
boolean | none | DDP: tolerate parameters without gradients | |
distributed.strategy |
string | Auto |
Auto, DDP, FSDP, ZeRO1, ZeRO2, ZeRO3 |
Auto and DDP replicate the model; FSDP and ZeRO shard it |
epochs |
number | 1 |
0.01 to 100 | Passes over the data when steps is -1 |
evalSteps |
integer | none | 1 to 100000 | Steps between validation passes; unset evaluates at the end only |
evaluationBatchSize |
integer | none | 1 to 1024 | Prompts generated together during Environment baseline and final evaluation |
generationBatchSize |
integer | none | 2 to 65536 | Total sampled completions per generation batch; must divide evenly across GPUs and numGenerations |
gradientAccumulation |
integer | 4 |
1 to 1024 | Steps whose gradients are summed before an update |
gradientCheckpointing |
boolean | true |
Recompute activations to save GPU memory | |
learningRate |
number | 0.000001 |
0 to 1 | Peak learning rate |
loggingSteps |
integer | 10 |
1 to 10000 | Steps between metric reports |
lora.alpha |
integer | 16 |
1 to 4096 | Adapter scaling numerator |
lora.dropout |
number | 0.05 |
0 to 0.9 | Adapter dropout |
lora.r |
integer | 8 |
1 to 1024 | Adapter rank |
lora.targetModules |
any | none | Module names, or all-linear | |
lrScheduler |
string | cosine |
cosine, linear, constant, constant_with_warmup |
Learning-rate schedule |
maxCompletionLength |
integer | 64 |
8 to 32768 | Tokens per completion; in a multi-turn episode, tokens per reply |
maxEpisodeTokens |
integer | none | 8 to 131072 | Tokens of a whole multi-turn episode after the prompt, the model’s and the Environment’s |
maxGradNorm |
number | none | 0 to 100 | Gradient clipping norm |
maxLength |
integer | 2048 |
16 to 131072 | Tokens per example after truncation (packed block size for pretraining) |
maxTurns |
integer | none | 1 to 31 | Model turns per episode of a multi-turn Environment; 1 grades the first reply alone |
method |
string | LoRA |
Full, LoRA, QLoRA |
Full fine-tuning, a LoRA adapter, or LoRA on a 4-bit quantized base (GPU only) |
numGenerations |
integer | 4 |
2 to 64 | Completions sampled per prompt (the GRPO group) |
precision |
string | Auto |
Auto, bf16, fp16, fp32 |
Numeric precision; Auto is bf16 where the GPU supports it, else fp16 |
sampling.temperature |
number | 0.7 |
0 to 5 | Sampling temperature |
sampling.topK |
integer | none | 0 to 1000000 | Keep the highest-probability K tokens; 0 disables top-k filtering |
sampling.topP |
number | 1 |
0 to 1 | Nucleus sampling mass |
saveSteps |
integer | 25 |
1 to 100000 | Steps between recovery checkpoints |
saveTotalLimit |
integer | none | 1 to 10 | Recovery checkpoints kept in the state directory |
seed |
integer | 42 |
0 to 2147483647 | Seed for data order, splits and initialization |
steps |
integer | 50 |
-1 to 1000000 | Optimizer steps; -1 trains for epochs instead |
taskFilter |
object | none | at most 8 keys | Keep only tasks whose metadata has these values (such as family on reasoning-gym), before trainTasks and heldOutTasks count them |
warmupRatio |
number | 0.03 |
0 to 0.5 | Fraction of steps spent warming up the learning rate |
weightDecay |
number | none | 0 to 1 | AdamW weight decay |
Outputs
Section titled “Outputs”| Output | Path |
|---|---|
outputs |
/nodus/outputs |
Every run also writes results.json, provenance.json and a sha256 manifest.json. Each file under /nodus/outputs is an output of the TrainingJob, named by its path: download one with nodus cp tj/<name>:outputs/results.json ./results.json.
Presets
Section titled “Presets”Without spec.resources, the TrainingJob takes the resources of the preset matching its model, method and quantization, else the runtime default.
| Model | Method | Quantization | GPUs |
|---|---|---|---|
| Qwen/Qwen3-0.6B | LoRA | none | 1 × RTX-4090 or L40S or RTX-A6000 |
| Qwen/Qwen3-1.7B | LoRA | none | 1 × RTX-4090 or L40S or RTX-A6000 |
| Qwen/Qwen3-4B | QLoRA | nf4 | 1 × RTX-4090 or L40S or RTX-A6000 |
| Qwen/Qwen3-4B | LoRA | none | 1 × L40S or RTX-A6000 or A100-80G |
| Qwen/Qwen3-8B | QLoRA | nf4 | 1 × L40S or RTX-A6000 or A100-80G |
Default resources: 1 × RTX-4090 or L40S or RTX-A6000 or A100-80G.
Estimates
Section titled “Estimates”The dry-run estimate multiplies these by the run’s steps or tasks; the nightly runs re-measure them.
| Accelerator | Seconds | Per |
|---|---|---|
| A100-80G | 0.7 | secondsPerStep |
| H100 | 0.45 | secondsPerStep |
| L40S | 1 | secondsPerStep |
| RTX-4090 | 1.2 | secondsPerStep |
| RTX-A6000 | 1.3 | secondsPerStep |
Example
Section titled “Example”apiVersion: nodus.dev/v1beta1kind: TrainingJobmetadata: {name: my-grpo-lora}spec: runtime: nodus/grpo-lora model: {uri: "hf://Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca"} environment: {name: nodus/graph-coloring@1.0.0, trainTasks: 200, heldOutTasks: 64, seed: 42} maxCostUSD: "5.00"