Skip to content

Training (Beta)

View Markdown

Beta: this feature may change before general availability

A TrainingJob says what to train (a base model, a dataset or an environment, and the runtime’s parameters) and Nodus runs it as a Job: it caches the model, picks GPUs from the runtime’s presets, estimates the run before it starts, checkpoints it, and collects the trained weights as outputs.

Terminal window
$ pip install nodus-compute
$ nodus login

A TrainingRuntime is a pinned trainer image with a parameter schema, presets and measured step times. The managed runtimes live in the nodus catalog project:

Runtime What it trains Data
nodus/sft Supervised fine-tuning, full or LoRA/QLoRA; task: Pretrain for continued or from-scratch pretraining prompt/completion or text JSONL
nodus/dpo Direct preference optimization prompt, chosen, rejected
nodus/orpo Odds-ratio preference optimization (no reference model) prompt, chosen, rejected
nodus/kto Kahneman-Tversky optimization from thumbs-up/down labels prompt, completion, label
nodus/reward-model A reward model for RLHF prompt, chosen, rejected
nodus/grpo-lora GRPO reinforcement learning against an Environment (single node) an Environment
nodus/distill Knowledge distillation from a teacher model prompts or chats
nodus/evaluate Evaluation only (mode: Evaluate) an Environment or a benchmark
Terminal window
$ nodus get trainingruntimes -n nodus
$ nodus get trainingruntime/sft -n nodus -o yaml # the parameter schema, presets and step times

In the console, Environments lists the same runtimes under Training templates. You start a run from the SDK or the CLI; the console shows it under Runs › Post-training with its stages, curves, results, cost and outputs.

sft.yaml
apiVersion: nodus.dev/v1beta1
kind: TrainingJob
metadata:
name: support-sft
spec:
runtime: nodus/sft
model:
uri: hf://Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca
secret: hf-token # only for gated repositories
data:
volume: support-chats # a Volume holding chats.jsonl
path: chats.jsonl
format: JSONL
columns: {prompt: question, completion: answer}
parameters:
steps: 2000
lora: {r: 16, alpha: 32}
maxCostUSD: "20.00"

Review it before anything runs. A dry run compiles the TrainingJob into the Job it will create and prices it:

Terminal window
$ nodus apply -f sft.yaml --dry-run -o yaml # status.estimate: cost to completion, first hold, start time
$ nodus apply -f sft.yaml

What the dry run decides:

  • Resources come from the first runtime preset that matches the model, the method (Full, LoRA or QLoRA, read from parameters) and quantization. Anything you set under spec.resources wins.
  • The estimate is the number of steps times the measured seconds per step on the slowest GPU type you allowed, plus evaluation tasks times seconds per task. Set spec.expectedDuration to use your own figure instead.
  • Parameters are merged over the runtime’s defaults and checked against its schema. A typo such as learning_rate is rejected with the field path, before you are charged anything.

Model, teacher, and dataset Volume inputs are fixed to their latest committed revision when you create the TrainingJob, unless you select an explicit revision. Upload the files and wait for the Volume to be Ready first. Later uploads do not change the admitted run.

Pin Hugging Face model revisions to a 40-character commit. The first TrainingJob that names a revision imports it once into a read-only Volume in your project, and every later run on that revision reuses it. The trainer runs with HF_HUB_OFFLINE=1, so billed GPU time is never spent downloading weights.

Terminal window
$ nodus get tj -w
NAME RUNTIME PHASE STAGE CHANGE COST AGE
support-sft nodus/sft Running Training $0.12 6m
$ nodus describe tj/support-sft
$ nodus logs job/support-sft -f

STAGE is CachingModel while the model imports, then the stage the trainer reports (Baseline, Training, Evaluating), and Exporting while the outputs are written to spec.exportTo. describe adds the step, total steps, loss and tokens per second, the provenance of the run (runtime and environment digests, model revision, and whether the runtime is a managed one) and the cost so far.

In the console, open the run from Compute: the Overview shows the stages, step, loss and the baseline and final pass rates; Gradings lists every graded task as it is scored; Results has the outputs and the before and after table.

The TrainingJob creates a Job of the same name and owns it. Logs, checkpoints and outputs are the Job’s, so everything in the Jobs guide applies.

Set spec.exportTo.volume to a Ready ReadWriteOnce Volume in the same project to also copy the complete result there. The runtime must declare an output at /nodus/outputs. The export uses an isolated CPU Job under the TrainingJob’s spending cap, waits up to 5 minutes for capacity, and allows 15 minutes for the copy. It follows suspension and cancellation. Success records the exact committed Volume revision in the Exported condition; files remain downloadable as Job outputs if the copy fails. Source or destination symbolic links are rejected.

When the run succeeds, every file the trainer wrote under /nodus/outputs is an output of the TrainingJob, named by its path: results.json (the final metrics), provenance.json, manifest.json (a sha256 per file) and the weights under adapter/ (LoRA and QLoRA) or model/ (full fine-tuning), with the tokenizer beside them.

Terminal window
$ nodus get tj/support-sft -o jsonpath='{.status.outputs[*].name}'
$ nodus cp tj/support-sft:outputs/results.json ./results.json
$ nodus cp tj/support-sft:outputs/adapter/adapter_model.safetensors ./adapter_model.safetensors

nodus cp verifies each file against the sha256 recorded when it was collected.

Leave out the lora block for full fine-tuning; presets for the Full method pick larger GPUs.

Pretraining uses the same runtime with task: Pretrain. initialization: Continued (the default) continues from the model’s weights; Scratch uses only its architecture and tokenizer:

spec:
runtime: nodus/sft
task: Pretrain
model: {uri: hf://Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca}
data: {volume: corpus, format: JSONL}
parameters: {steps: 10000, initialization: Scratch, packing: true}
resources: {gpu: {type: [h100-sxm-80g], count: 8}}

dpo, orpo, kto and reward-model take a dataset whose columns you map onto the names the runtime expects:

spec:
runtime: nodus/dpo
model: {volume: {name: my-sft-model, revision: 3}} # a model you trained earlier, by Volume
data:
volume: prefs
format: JSONL
columns: {prompt: prompt, chosen: chosen, rejected: rejected}
parameters: {steps: 400}

A model can come from a Hugging Face URI or from one of your Volumes at a pinned revision, so a DPO run can start from the output of an SFT run.

distill trains the student in model on the teacher’s distributions. The teacher is cached the same way:

spec:
runtime: nodus/distill
model: {uri: hf://Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca}
teacher: {uri: hf://Qwen/Qwen3-4B@0e9e39f249a16976918f6564b8830bc894c89659}
data: {volume: chats}
parameters: {temperature: 2.0}
resources: {gpu: {type: [h100-sxm-80g], count: 2}}

Reinforcement learning (Beta, single node)

Section titled “Reinforcement learning (Beta, single node)”

grpo-lora trains against an Environment: it samples completions for the environment’s tasks and Nodus grades them. The trainer never sees the expected answers and never grades itself: it submits completions to the TrainingJob’s gradings endpoint, and Nodus runs the environment’s grader in isolated containers on the same GPU host, with no network access, model files or GPU devices and the answers readable only by the grader. Task generation also uses this host; Nodus does not provision a separate CPU machine. A grader that fails is reported as an infrastructure failure and never scores the completion 0.

Environment-backed training and evaluation require GPU capacity that supports isolated grading. Grading is serialized per TrainingJob. Baseline and final model evaluation run on the GPU. Custom runtime images must implement the task-fetch endpoint and declare spec.environmentTasks: AttemptAPI; older runtimes are rejected before GPU acquisition.

spec:
runtime: nodus/grpo-lora
model: {uri: hf://Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca}
environment: {name: nodus/graph-coloring@1.0.0, trainTasks: 50, heldOutTasks: 64, seed: 42}
evaluation: {baseline: true, final: true}
parameters: {steps: 50, lora: {r: 8}}
resources: {gpu: {type: [rtx4090-24g], count: 1}}
maxCostUSD: "5.00"

With evaluation.baseline and evaluation.final, the run scores the held-out tasks before and after training. status.summary holds both pass rates and a comparison: the change in percentage points, and whether it is comparable at all. A comparison needs the same task set on both sides, at least 16 scored tasks, and a result that is not all-pass or all-fail on both sides; otherwise its reason says why (TooFewTasks, DifferentTasks, NonePassed, AllPassed, IncompleteEvents). status.diagnostics.rewardSignal warns when every training reward was the same, which teaches a group-relative method nothing.

Reproducing the legacy letter-counting protocol

Section titled “Reproducing the legacy letter-counting protocol”

Use nodus/letter-counting-legacy-eval@1.0.0 for the legacy evaluation protocol: the first 64 tasks from Reasoning Gym’s letter_counting generator with seed 42, an answer-only system message, greedy decoding with 256 completion tokens, and the generator’s raw-completion scorer. The letter-counting catalog example pins Qwen3-1.7B and generates one prompt at a time.

Use nodus/letter-counting-legacy-rl@1.0.0 for the legacy RL protocol. Its first 64 tasks are training data; the following 64 are held out for both baseline and final evaluation. The example sets 50 optimizer steps, learning rate 5e-6, eight generations, generation batch size eight, per-GPU batch size one, gradient accumulation four, temperature 1.2, top-p 0.95, top-k 50, and 64 completion tokens. It uses a rank-eight LoRA adapter with alpha 16, no dropout, all linear layers, bf16, non-reentrant gradient checkpointing, and a linear learning-rate schedule without warmup. Thinking is disabled in both phases. Recovery checkpoints remain enabled.

These profiles pin Reasoning Gym 0.1.25 and preserve the legacy source’s task slicing, prompt roles and scoring. They do not establish historical binary equivalence or a measured pass rate. The evaluation profile’s 64 tasks are different from the RL profile’s held-out 64 tasks; compare baseline and final within the same RL run. The existing nodus/reasoning-gym profile uses shuffled family-specific seeds and remains a different benchmark. Record the exact runtime, environment and model digests alongside results before claiming a numeric comparison.

RL runs on one node. Multi-node RL and custom environment collections are not available yet.

Set resources.nodes above 1 to train across several machines, each with resources.gpu.count GPUs. The TrainingJob compiles into one multi-node Job (a gang): every node starts together, torchrun is launched on each with its rank and the rendezvous address, and the run checkpoints with torch.distributed.checkpoint.

spec:
runtime: nodus/dpo
model: {volume: {name: my-8b, revision: 3}}
data: {volume: prefs, format: JSONL, columns: {prompt: prompt, chosen: chosen, rejected: rejected}}
parameters: {steps: 400}
resources: {gpu: {type: [h100-sxm-80g], count: 8}, nodes: 2}
distributed: {network: Global, transport: Auto}
  • Launchers. A runtime declares Torchrun, Accelerate, Deepspeed or Plain, and the run uses that launcher on every node: torchrun reads the gang’s rendezvous settings, accelerate launch and DeepSpeed’s per-node launcher read the rank, address and world size Nodus sets on each node, and Plain starts the runtime’s command once per node. A single node runs torchrun --standalone for a Torchrun runtime, accelerate launch or deepspeed for the other two.
  • Placement. distributed.network is Colocated (one provider and region, the fastest interconnect), Regional or Global; transport: Auto also allows an encrypted SSH connection between two restricted GPU machines without a separate relay server. Native GPU-to-GPU connectivity is preferred. Wider settings find capacity sooner at some cost in throughput.
  • Recovery. If a node is lost, the whole gang restarts from the latest checkpoint every rank committed, up to 3 times; progress since that checkpoint is redone. A single-node run restores its latest snapshot instead. Managed trainers bind recovery state to the full run configuration, runtime image digest, and pinned model and dataset references; incompatible state is refused. GRPO reuses the baseline saved in that state so recovery does not repeat baseline grading.
  • Cost. The estimate covers every node. Assembly time (nodes waiting for each other) is billed and bounded; see Multi-node training for the limits.

spec.maxCostUSD caps the whole TrainingJob: its Job, its same-host grading, its output export and its retries. When the cap is reached the run is suspended, not failed, and active grader containers are removed. Raise the cap to resume the same run from its checkpoint:

Terminal window
$ nodus patch tj/support-sft --type merge -p '{"spec":{"maxCostUSD":"40"}}'

The cap can only be raised. spec.state: Suspended pauses a run yourself, and Running resumes it.

spec.tracking.mlflow (uri, secret, experiment) or spec.tracking.wandb (connection) send the trainer’s metrics to your own MLflow server or Weights & Biases project as well. Nodus shows its own stage and metrics either way. The catalog runtimes report to MLflow when the run has MLFLOW_TRACKING_URI and to W&B when it has WANDB_API_KEY, which is what those two fields set; the Secret holds MLFLOW_TRACKING_TOKEN (or MLFLOW_TRACKING_USERNAME and MLFLOW_TRACKING_PASSWORD), and the W&B Connection’s Secret holds WANDB_API_KEY.

A TrainingRuntime in your project can run any image. The trainer reads:

Variable Contents
NODUS_INPUT_PARAMS a directory that holds params.json, the run: task and mode (for example SFT, Train), the merged, schema-checked parameters, data (path, format and column mapping), environment (task counts, seed and graders), evaluation, trainingJob and the pinned modelRevision and teacherRevision
Environment tasks Call POST …/trainingjobs/{name}/tasks through the attempt API socket. The response files maps train.jsonl and test.jsonl to JSONL strings containing {taskId, prompt, metadata} and no hidden answers. Managed runtimes fetch them automatically; custom runtimes must implement this step.
NODUS_INPUT_MODEL, NODUS_INPUT_TEACHER, NODUS_INPUT_DATA where the model, teacher and data are mounted
NODUS_STATE_DIR, NODUS_OUTPUT_DIR resumable state and results
NODUS_NODE_RANK, MASTER_ADDR, PET_* the gang of a multi-node run
NODUS_CHECKPOINT_URI, NODUS_RESTORE_URI where ranks write and restore distributed checkpoints

Write and load your own model, optimizer and progress files under the state directory; Nodus decides when to checkpoint and where checkpoints are stored, and never adds resume flags to your command.