# Models and agents

---

Hosted inference, durable agents and model training (Beta).

---

# Agents

> Run an agent on Claude in its own sandbox, send it follow-up prompts, and pay for the model from your credits.

Source: https://nodus-platform-site.pages.dev/docs/guides/agents/
Build revision: 211ad9f836655b1c3a2668c4693e442471f28614

**Beta:** this feature may change.

An **Agent** is a definition: a system prompt, the Claude access it runs on and a cost cap. An **AgentRun** is one conversation with it. Each run works in its own sandbox, where the model runs shell commands to check its work, and it survives restarts of Nodus by replaying recorded results. An interrupted command or model call using your own key parks with `NeedsResolution` when its outcome is uncertain; inspect its effects before starting replacement work.

Agents, parallel AgentGroups and Environment evaluation batches are in Beta. Each run is one agent conversation.

## Sign in and create an agent

agent.yaml

```yaml
apiVersion: nodus.dev/v1
kind: Agent
metadata:
  name: hello-agent
spec:
  system: |
    You are a careful assistant. Run shell commands in your sandbox to check your work, and answer in one short
    paragraph.
  # Nodus's Claude: each prompt is routed to Haiku, Sonnet, Opus or Fable, and billed from your credits.
  model:
    access: Nodus
  perRunMaxCostUSD: "0.50"
```

Terminal window

```bash
nodus login
nodus apply -f agent.yaml
```

The agent uses Nodus’s Claude. Every prompt goes to a small router that picks the right Claude model for it (Haiku for simple requests up to Fable for demanding ones). That model answers the whole prompt, including its commands, so the model never changes in the middle of a turn. If the router is unavailable, Sonnet answers. Restrict the choice with `spec.model.families`, for example `[haiku, sonnet]`.

## Start a run

run.yaml

```yaml
apiVersion: nodus.dev/v1
kind: AgentRun
metadata:
  name: hello-agent-run
spec:
  agent: hello-agent
  input: "Use the shell to print the Python version, then tell me which one it is."
```

Terminal window

```bash
nodus apply -f run.yaml
nodus wait agentrun/hello-agent-run --for=jsonpath='{.status.phase}'=Succeeded --timeout=8m
nodus get agentrun/hello-agent-run -o yaml
```

Or without a file: `nodus create agentrun triage --agent hello-agent --prompt "Summarize the logs"`. To start from Nodus’s ready-made assistant, name the template `claude-assistant`.

`status.turns` lists each prompt with the model that served it, how it was chosen (`jev` or `fallback`) and its cost. `status.answer.preview` holds the first 4 KiB of the final answer; `nodus agentrun answer` prints all of it.

Terminal window

```bash
nodus agentrun answer hello-agent-run
nodus agentrun steps hello-agent-run
```

`steps` lists what the run did in order: creating its sandbox, each routing and model call, each command. A step that shows `Completed` is never repeated, even after a restart.

The console’s Agents page shows the same: each run with the model of its latest turn and its cost, and on a run every turn with the model that served it, how that model was chosen and what the turn cost, then the answer and the recorded steps.

## From Python

```python
import nodus

agent = nodus.ClaudeAgent("helper", system="You are careful.", families=["haiku", "sonnet"])
print(agent.remote("Use the shell to print the Python version"))   # runs to the end, returns the answer

run = agent.submit("Summarize the logs", keep_alive=True)          # a handle: run.answer(), run.steps(), run.cancel()
run.send("message", "Now the staging logs")
```

`ClaudeAgent` creates the agent on first use. Pass `api_key_secret="anthropic-key"` to use your own Anthropic key, or `ClaudeAgent.from_name("claude-assistant", project="nodus")` to run the ready-made template. To run many at once from Python, see [Run agents in parallel](https://nodus-platform-site.pages.dev/docs/guides/python/agents/#run-agents-in-parallel).

## Keep a run open for follow-ups

Set `spec.keepAlive: true` and the run waits for messages after each turn instead of finishing. While it waits it uses no model and keeps its sandbox, which can still accrue compute charges. Set `spec.deadline` to bound the run’s lifetime. The deadline also applies while waiting for a message: the run fails with `DeadlineExceeded` and its sandbox is deleted without another model call.

Terminal window

```bash
nodus apply -f run.yaml                                   # with keepAlive: true
nodus agentrun send hello-agent-run "Now do the same for the staging logs" --key staging-1
```

Sending the same `--key` again delivers the message once. A run that is still working reads the message after its current turn; a finished run refuses it with `AgentRunFinished`. Cancel a run with `nodus cancel agentrun/NAME`; its sandbox is deleted.

The same calls are REST: `POST …/agentruns/NAME/messages` with `{"payload": "…", "messageKey": "…"}`, and `GET` on `…/steps` and `…/answer`.

## Run agents in parallel

An **AgentGroup** runs many runs of one agent at once. You choose how many run at the same time and how much the whole group may spend. The other runs wait their turn, and one command cancels or deletes all of them. Create the agent and the group, then one run for each task:

agent.yaml

```yaml
apiVersion: nodus.dev/v1
kind: Agent
metadata:
  name: parallel-agents-worker
spec:
  system: |
    Follow the task's requested answer format; otherwise answer in one short sentence.
  # Nodus's Claude, limited to Haiku so each run stays far below its cap.
  model:
    access: Nodus
    families: [haiku]
  maxTokens: 1024
  perRunMaxCostUSD: "0.10"
```

group.yaml

```yaml
apiVersion: nodus.dev/v1
kind: AgentGroup
metadata:
  name: parallel-agents
spec:
  agent: parallel-agents-worker
  limits:
    maxActive: 2          # at most two runs go at once
  maxCostUSD: "0.60"      # a run starts only while its own cap still fits under what is left
```

runs.yaml

```yaml
apiVersion: nodus.dev/v1
kind: AgentRun
metadata:
  name: parallel-agents-a
spec:
  agent: parallel-agents-worker
  group: parallel-agents
  taskKey: a
  input: "What is the capital of France?"
---
apiVersion: nodus.dev/v1
kind: AgentRun
metadata:
  name: parallel-agents-b
spec:
  agent: parallel-agents-worker
  group: parallel-agents
  taskKey: b
  input: "What is the capital of Japan?"
---
apiVersion: nodus.dev/v1
kind: AgentRun
metadata:
  name: parallel-agents-c
spec:
  agent: parallel-agents-worker
  group: parallel-agents
  taskKey: c
  dependsOn: [a, b]       # c waits until a and b have succeeded
  input: "Say in one sentence that both questions have been answered."
```

Terminal window

```bash
nodus apply -f agent.yaml
nodus apply -f group.yaml
nodus apply -f runs.yaml
nodus patch ag/parallel-agents --patch '{"spec":{"sealed":true}}'
```

Or without a file: `nodus create agentgroup parallel-agents --agent parallel-agents-worker --max-active 2 --max-cost 0.60`.

Each run is a task of the group: `spec.group` names the group and `spec.taskKey` names the task (a DNS label that is unique in the group). The run is named `<group>-<taskKey>`; leave `metadata.name` out or set exactly that. The run uses the group’s agent, so `spec.agent` must match it. The last command seals the group, which says that no more runs are coming (see below).

### Watch progress and read the answers

Terminal window

```bash
nodus wait ag/parallel-agents --for=jsonpath='{.status.phase}'=Succeeded --timeout=12m
nodus get ag
nodus get ar --field-selector spec.group=parallel-agents
nodus agentrun answer parallel-agents-c
```

`nodus get ag` lists each group with its phase, `RUNS` (finished successfully out of all) and `ACTIVE` runs, and its cost. `status.counts` splits the runs into `queued`, `active`, `waiting` (for a message or for funds), `succeeded`, `failed` and `cancelled`. `status.blockedReasons` says why queued runs have not started and how many wait for each reason: `DependencyWait`, `MaxActiveReached` or `MaxCostReached`. A queued run’s own `status.reason` is `DependencyWait`, or `GroupWait` while it waits for a free slot or for room under the cost cap. The runs you list are ordinary runs, so `status.answer`, `nodus agentrun answer` and `nodus agentrun steps` work on each of them.

A group is `Running` until it is sealed and every run has finished. It then becomes `Succeeded`, or `Failed` with reason `RunsFailed` when a run failed or was cancelled because a dependency failed. A sealed group with no runs succeeds.

### How many run at once

`spec.limits.maxActive` is how many runs the group lets run at the same time (default 10, 1 to 1000). Change it on a live group:

Terminal window

```bash
nodus patch ag/parallel-agents --patch '{"spec":{"limits":{"maxActive":4}}}'
```

`spec.limits.maxPending` is how many unfinished runs the group takes in (default 10000); a run beyond it is refused.

### Dependencies

`spec.dependsOn` lists the task keys a run waits for. It starts only after every one of them has succeeded, and until then it is queued with reason `DependencyWait`. A dependency must already be in the group, so create the runs it depends on first (earlier in the same file works). When a dependency fails, the run is cancelled with reason `DependencyFailed`; when a dependency was cancelled, the run is cancelled with reason `DependencyCancelled`.

### The group cost cap

`spec.maxCostUSD` is one cap for the whole group. A run starts only when its own cap (the agent’s `perRunMaxCostUSD`) still fits under what is left: the cap minus what the group has spent and minus the caps of the runs that have started and not finished. The members together can therefore never spend more than the cap, and runs that do not fit yet wait with `MaxCostReached`. In the example, the $0.60 cap has room for the caps of six runs at $0.10 each before any run has spent anything. Raise the cap on a live group with `nodus patch`; it cannot be lowered. If the cap is smaller than one run’s cap, no run starts: raise it or lower the agent’s `perRunMaxCostUSD`. `status.cost` shows `totalUSD` and `limitUSD`. The cap covers the runs’ Claude usage (model and routing calls), the spend each run’s `perRunMaxCostUSD` limits; the sandbox each run works in is billed on its own and is not counted against either cap.

### Seal, cancel and delete

A group takes new runs until you seal it with `spec.sealed: true`. Sealing cannot be undone, and a sealed group refuses new runs with a conflict. A group finishes only once it is sealed, so seal it after you create the last run.

Terminal window

```bash
nodus cancel ag/parallel-agents
nodus delete ag/parallel-agents
```

`nodus cancel` cancels every run that has not finished and leaves the finished runs, with their answers, as they are. `status.phase` goes through `Cancelling` to `Cancelled`. `nodus delete` removes the group together with all of its runs.

## Evaluate an agent

Create an evaluation group with an Environment at an explicit version. Nodus generates the fixed task batch, runs it through ordinary AgentRuns, and grades completed answers using the Environment’s isolated grader. There is no manual submit or seal step. Set a project Budget to limit total spend: the group cap covers model usage, while agent and grading Sandbox compute is billed separately.

After deploying the example’s `parallel-agents-worker` Agent above, run this Python example. It creates two held-out tasks, waits for a terminal result, prints per-case outcomes and deletes the evaluation group afterward.

evaluation.py

```python
"""Evaluate the deployed example agent on a fixed, versioned task batch."""

import nodus


def main() -> None:
    # Model caps exclude Sandbox compute; set a project Budget before running this example.
    evaluation = nodus.AgentGroup.create(
        "parallel-agents-eval",
        agent="parallel-agents-worker",
        max_active=2,
        max_cost="0.20",
        evaluation={
            "environment": "nodus/arithmetic-v2@2.0.0",
            "split": "test",
            "tasks": 2,
            "repetitions": 1,
            "seed": 42,
            "timeout": "10m",
        },
    )
    try:
        status = evaluation.wait(timeout=660)
        results = evaluation.results()
        print(
            f"{status.phase}: cases={len(results)}, pending={status.evaluation.pending}"
        )
        print("pass rate:", status.evaluation.get("passRate"))
        print("mean reward:", status.evaluation.get("meanReward"))
        for result in results:
            print(result.task_id, result.state, result.get("verdict"))
    finally:
        evaluation.delete()


if __name__ == "__main__":
    main()
```

The batch is at most 100 cases including repetitions. Every member pins the same Agent revision; results record the Environment version, resolved image digest, split and seed. Only task prompts reach the agent, never hidden answers. The timeout defaults to 30 minutes, may be 1 minute to 24 hours, and includes manifest preparation, queueing and funding waits from group creation. Each member receives that absolute deadline. Grading also stops at the deadline; completed answers that could not be scored are shown as evaluation timeouts, excluded from pass rate and mean reward.

Open the group in the console to see scored cases, pass rate, mean reward, pending cases and execution/grading failures separately. The Results view links each case to its AgentRun and shows grader evidence. A failed run, missing answer or grader failure is unscored and does not become a zero reward. A wrong or invalid answer is a scored result. Per-case records share the member journal’s 30-day payload retention; deleting the group removes its runs and records. Canceling or finishing the group cleans up its grading pool.

## Use your own Anthropic key

Store the key in a Secret under the key `ANTHROPIC_API_KEY` and point the agent at it. Nodus then charges nothing for the model, and your Anthropic account pays for it.

```yaml
spec:
  model:
    access: BYOK
    name: claude-sonnet-5-5
    apiKeySecret: anthropic-key
```

## What a run costs

On Nodus’s Claude each model call, and the routing call before each turn, is billed from your credits at cost divided by 0.875. `spec.perRunMaxCostUSD` (default $1.00) caps what a run spends on the model. Before each call the run checks that the call, at its largest (the whole conversation as input and `spec.maxTokens` of output), still fits under the cap. When it does not, the run waits (`reason: MaxCostReached`) and makes no more calls. The cap is copied from the agent when the run is created, so changing the agent does not change a run that has started: to continue the work, cancel the run and start a new one. If your credits run out, the run waits instead and continues when you add credits. A run on your own Anthropic key is not charged by Nodus and has no cap to reach. Sandbox compute is billed as for any sandbox. `status.cost` shows the model spend so far. Runs in a group also share the group’s cap (see [Run agents in parallel](https://nodus-platform-site.pages.dev/docs/guides/agents/#run-agents-in-parallel)).

## Limits

* One run works one prompt at a time, with at most `spec.maxToolCalls` commands per turn (default 50).
* The input and each message are at most 256 KiB. A run’s recorded payloads are kept 30 days after it ends.
* A run created for a stopped agent (`spec.state: Stopped`) is refused with `AgentStopped`.
* A group takes at most 10000 unfinished runs. Create runs in batches of up to 100 by posting an `AgentRunList` to `…/agentruns` with an `Idempotency-Key` header; a batch is all-or-nothing, and its runs are created in list order, so list each run after the runs it depends on.

See [Durable execution](https://nodus-platform-site.pages.dev/docs/concepts/durable-execution/) for what survives a restart.


---

# Environments

> Task sets with graders for reinforcement learning and evaluation, from the Nodus catalog or your own image.

Source: https://nodus-platform-site.pages.dev/docs/guides/environments/
Build revision: 211ad9f836655b1c3a2668c4693e442471f28614

An Environment is a versioned set of tasks with a grader: a prompt for each task, and a program that decides whether a completion is correct. TrainingJobs use Environments for reinforcement learning and evaluation, and agent evaluations use the same ones. Environments are generally available (`nodus.dev/v1`).

## Browse the catalog

Terminal window

```console
$ nodus get environments -n nodus
NAME               VERSION   CATEGORY    MODES             PHASE
graph-coloring     1.0.0     Reasoning   Train, Evaluate   Ready
arithmetic         2.0.0     Math        Train, Evaluate   Ready
gsm8k              1.0.0     Math        Train, Evaluate   Ready
reasoning-gym      1.0.0     Reasoning   Train, Evaluate   Ready
python-functions   1.0.0     Code        Train, Evaluate   Ready
$ nodus describe environment/graph-coloring -n nodus
```

`describe` shows the summary, licenses, the size of each split, a sample task, the graders and the examples: ready TrainingJob templates with measured results. The console’s Training page lists the same catalog, and “Run example” starts a TrainingJob from one; in Python, `nodus.TrainingJob.from_example("nodus/graph-coloring@1.0.0")` does the same.

## Use one in a TrainingJob

Name the environment and version, and how many tasks to train and evaluate on:

```yaml
spec:
  runtime: nodus/grpo-lora
  environment:
    name: nodus/graph-coloring@1.0.0
    trainTasks: 50       # from the train split
    heldOutTasks: 64     # from the test split, never trained on
    seed: 42
```

The same name, version and seed always give the same tasks in the same order, so two runs are comparable. The train and test splits never share a task.

## How grading works

* Tasks carry a prompt and non-secret metadata only. Expected answers stay with the grader; the trainer, the model and your code never see them.
* Completions are graded by Nodus, not by the trainer. Program graders run in grading Sandboxes that Nodus creates for your TrainingJob or agent evaluation: no network access, deleted after five minutes idle, at most `grading.maxParallel` at once. Code from a completion runs as a separate user that cannot read the answers. `ExactMatch` graders compare the completion with the answer inside Nodus, without a Sandbox.
* Every verdict is `Correct`, `Incorrect`, `InvalidOutput` or `InfrastructureFailure`, with a reward and evidence recorded by the host that ran the grader: the process, its exit code, its duration and a hash of its output.
* `InvalidOutput` is a completion the grader cannot parse, such as an answer without the expected tag. It is a failed task with a reward of 0.
* `InfrastructureFailure` means the grader could not run. It carries no reward and does not count toward a pass rate, so an outage never looks like a wrong answer.
* Grading Sandboxes are billed to the TrainingJob and count toward its `maxCostUSD`. They are deleted when the run finishes or is suspended.

## Train on a reward function in one call

When your tasks and reward are in Python, `rl.train` is all you need. To see it work first, run the example that ships with the SDK:

Terminal window

```sh
python -m nodus.examples.rl
```

It is one file, `nodus/examples/rl.py`, and it is the template for your own task:

```python
import re
from nodus.recipes import rl

ANSWER = re.compile(r"<answer>\s*([A-Za-z]+)\s*</answer>")
tasks = [(f"Spell the word backwards inside <answer></answer>.\n\nWord: {w}", w[::-1]) for w in WORDS]

def reward(completion, answer):
    found = ANSWER.findall(completion)
    return None if not found else float(found[-1].lower() == answer)

run = rl.train(tasks, reward, max_cost=2)
run.watch()                          # each stage, then the reward, loss and KL of every step
run.outputs.download("./outputs")    # the LoRA adapter and the before-and-after comparison
```

`run.tasks(phase="Evaluation", outcome="Failed")` lists the graded tasks the run has scored so far, each with its `taskId`, `reward` and outcome, so you can read which tasks the trained model still gets wrong.

* **Tasks** are `(prompt, answer)` pairs or `{"prompt", "answer", "metadata"}` dicts. Only `reward` sees the answer.
* **Model.** The base model is `Qwen/Qwen3-0.6B` unless you pass `model=`, for example `model="Qwen/Qwen3-4B"`.
* **Held-out tasks.** A fifth of the tasks, and at least 16, is held out. The model is graded on them before and after training. Pass `test_tasks=` to choose them yourself.
* **What gets packaged.** The reward goes with the lines of its own file that it uses: imports, constants and helper functions. A name it cannot take along fails before anything is built.
* **Checked before it runs.** `rl.train` grades one task’s answer and an empty reply on your machine first. A reward that raises, or returns something other than a number, a bool or `None`, fails there instead of on a GPU. Return `None` or `0` when a reply holds no answer.
* **Reuse.** The same code and tasks reuse the same Environment, so a second run starts at once.
* **No network.** The reward runs in Nodus’s grader, which has no network access. Pass `pip=["package==1.2.3"]` for packages it imports.
* **How it learns.** Each step samples `group_size=8` replies to each of `groups_per_step=4` tasks and scores every reply against the others in its group. Replies are capped at `max_tokens=256`, and the learning rate is `learning_rate=4e-5`. You can change any of them.
* **Format rewards.** To reward the shape of an answer as well as its value, fold both into the reward, for example `return float(correct) - 0.1 * (not well_formed)`.
* **Other options.** `steps`, `gpu`, `max_cost`, `lora` and the other `grpo_lora` parameters are optional. The run appears under **Runs › Post-training** in the console, with its curves, results and cost.

To train on a catalog Environment, pass its name instead of tasks and a reward:

```python
run = rl.train("nodus/gsm8k@1.0.0", max_cost=5)
```

## Bring your own environment

Your own tasks and reward are a Python module with two functions. `tasks(split)` lists the prompts of the `train` or `test` split with their answers, and `reward(completion, answer)` scores one completion:

```python
def tasks(split):
    return [{"prompt": "Spell the word backwards...\n\nWord: valley", "answer": "yellav"}, ...]

def reward(completion, answer):
    found = ANSWER.findall(completion)
    if not found:
        return None                      # no answer in the completion: InvalidOutput
    return 1.0 if found[-1].lower() == answer else 0.0
```

A reward of 1 (or `True`) is `Correct`, anything lower `Incorrect`, and `None` `InvalidOutput`. If `reward` raises, the verdict is `InfrastructureFailure`, never a reward of 0. Answers reach only `reward`; the trainer and the model see the prompt and the optional `metadata`. A test prompt never appears in train, even if your lists repeat it.

The quickest way in needs no Docker: publish the module as a pip package whose `load_environment()` returns an object (or the module itself) with `tasks` and `reward`, and name the exact version. Nodus builds it into an image when you apply the Environment; if the package does not load, the build fails and the Image `env-<name>-<version>` shows the step and its log:

```yaml
apiVersion: nodus.dev/v1
kind: Environment
metadata:
  name: reverse-words
spec:
  version: 1.0.0
  package:
    pip: {name: reverse-words, version: 1.0.0}   # index: for a private index
    loader: load_environment                     # the default; or module:function
  category: Custom
  modes: [Train, Evaluate]
  rewardType: Binary
```

## Run an environment from the Environments Hub

An environment published on [Prime Intellect’s Environments Hub](https://app.primeintellect.ai/dashboard/environments) is a pip package, so it runs on Nodus as it is. Name it with the Hub’s index for its owner and the version the Hub lists:

```yaml
apiVersion: nodus.dev/v1
kind: Environment
metadata:
  name: reverse-text
spec:
  version: 0.1.4
  package:
    pip:
      name: reverse-text
      version: 0.1.4
      index: https://hub.primeintellect.ai/primeintellect/simple/
  category: Custom
  modes: [Train, Evaluate]
```

Its `load_environment()` returns a `verifiers` environment, and Nodus trains on it directly: the `dataset` is the train split, the `eval_dataset` the held-out one (with only one of them, every fifth task is held out), and each completion is scored by the environment’s own rubric, with `verifiers` 0.1 through 0.3. A multi-turn environment, such as a game like `wordle`, runs too: see [Train on a multi-turn environment](https://nodus-platform-site.pages.dev/docs/guides/environments/#train-on-a-multi-turn-environment), and so does one where the model calls tools: see [Train on a tool-calling environment](https://nodus-platform-site.pages.dev/docs/guides/environments/#train-on-a-tool-calling-environment). An environment whose tools run in a remote sandbox, a browser or an MCP server fails the build with that reason, as does one scored by another model (a `JudgeRubric`), since grading has no network and holds no credential for that model. The console marks each kind in its Hub search from the Hub’s own tags. The package comes from the Hub index alone and its dependencies may also come from PyPI.

Grading has no network. Nodus fetches the datasets an environment loads while it builds the image, and every task list and grade after that reads that copy, so a run always sees the data its build saw.

## Train on a multi-turn environment

In a multi-turn environment the model and the environment take turns: the model replies, the environment answers with feedback, and the episode goes on until the environment ends it. Set `maxTurns` on the TrainingJob (**Turns per episode** under **Advanced** in the console) to let an episode run that many model turns:

```yaml
spec:
  runtime: nodus/grpo-lora
  environment: {name: wordle@0.1.7, trainTasks: 512, heldOutTasks: 64, seed: 42}
  parameters:
    maxTurns: 6              # model turns per episode
    maxCompletionLength: 256 # tokens per reply
    maxEpisodeTokens: 2048   # tokens of the whole episode after the prompt, the model's and the environment's
```

The reward is the one the environment gives the whole episode, and training learns only from the model’s own tokens. An episode that reaches either cap ends there and is graded as it stands. With the default `maxTurns: 1`, only the first reply is graded. Baseline and final evaluation play the same episodes greedily.

An [OpenEnv](https://github.com/meta-pytorch/OpenEnv) environment runs in process from a package whose `load_environment()` returns it. Each task is a seed: Nodus resets the environment with it, and every action is the model’s reply as the one text field of the environment’s action class. The environment must play the same episode for the same seed. Seeds 0 to 799 train and 800 to 999 are held out, unless the environment sets `train_seeds` and `test_seeds`:

```python
import nltk
from textarena_env.server.environment import TextArenaEnvironment

def load_environment():
    try:  # the build downloads NLTK's word lists; grading reads them offline
        nltk.data.find("corpora/words")
        nltk.data.find("taggers/averaged_perceptron_tagger_eng")
        cached = True
    except LookupError:
        cached = False
    return TextArenaEnvironment("Wordle-v0", download_nltk=not cached)
```

A module of your own can also be multi-turn: give it `step(turns, answer)` in place of `reward`. It gets every model turn so far and returns the environment’s next message, or `None` once the episode ended, and the reward so far. Grading keeps no state between turns, so `step` replays the turns from the start.

## Train on a tool-calling environment

In a `verifiers` `ToolEnv` the model calls the environment’s tools and reads their results. Each task’s prompt offers the tools through the model’s chat template, the model calls one by writing a `<tool_call>{"name": "...", "arguments": {...}}</tool_call>` block (the form Qwen and most open chat templates teach), and the environment runs the call in its own code. The results are the next turn, and the episode ends when the model answers without calling a tool. Set `maxTurns` to the rounds of calls an episode may make:

```yaml
spec:
  runtime: nodus/grpo-lora
  environment: {name: tool-test@0.1.1, trainTasks: 21, heldOutTasks: 43, seed: 42}
  parameters:
    maxTurns: 3
    maxCompletionLength: 256
```

The tools run inside the grader, which has no network, so a tool that searches the web or runs code in a remote sandbox cannot train. Pick a model whose chat template supports tools.

## Run any Reasoning Gym family

The catalog’s `nodus/reasoning-gym` serves five reviewed families. To train on any other [Reasoning Gym](https://github.com/open-thought/reasoning-gym) family, or a mix of them, publish a package whose `load_environment()` returns the dataset; its own `score_answer` scores each completion:

```python
import reasoning_gym

def load_environment():
    return reasoning_gym.create_dataset("knights_knaves", size=2000, seed=7)
```

Every fifth entry is held out. The model’s last `<answer>…</answer>` is its answer, or the whole completion without one; partial credit is the reward and only a full score is `Correct`. Each family’s own licence applies.

## Run a verl or SkyRL dataset

A dataset prepared for [verl](https://github.com/volcengine/verl) or [SkyRL](https://github.com/NovaSky-AI/SkyRL) runs with its reward unchanged. Publish a module that names the parquet files with verl’s own keys, and your verl reward function under its usual name:

```python
train_files = "hf://datasets/BytedTsinghua-SIA/DAPO-Math-17k/data/dapo-math-17k.parquet"
val_files = "hf://datasets/BytedTsinghua-SIA/AIME-2024/data/aime-2024.parquet"   # optional

def compute_score(data_source, solution_str, ground_truth, extra_info=None):
    ...                                  # verl's custom reward signature; a dict's "score" also works
```

Each row is verl’s: `prompt` (a system message and one user message at most), `data_source`, `reward_model.ground_truth` and `extra_info`. Without `compute_score`, each row’s `env_class` names the SkyRL-gym environment that scores it, built from the row as SkyRL builds it; add `skyrl-gym` to the package’s dependencies. Without `val_files`, every fifth row is held out. The package needs `datasets` as a dependency, and a SkyRL environment of more than one turn fails the build with that reason.

To ship your own system packages or files, build the module into an image on the `env-base` image instead, push it and name its digest:

```dockerfile
FROM ghcr.io/nodus-compute/catalog/env-base:1.0.0
COPY reverse_words.py /opt/environment/
ENV NODUS_ENVIRONMENT_LOADER=reverse_words NODUS_ENVIRONMENT=reverse-words NODUS_ENVIRONMENT_VERSION=1.0.0
RUN nodus-env info
ENV HF_HUB_OFFLINE=1 HF_DATASETS_OFFLINE=1
```

```yaml
apiVersion: nodus.dev/v1
kind: Environment
metadata:
  name: reverse-words
spec:
  version: 1.0.0
  package: {image: registry.example.com/acme/reverse-words@sha256:...}
  category: Custom
  modes: [Train, Evaluate]
  rewardType: Binary
```

Terminal window

```console
$ nodus apply -f environment.yaml
$ nodus get environment/reverse-words -w      # Ready once the image digest is verified
```

A TrainingJob names it without the `nodus/` prefix (`environment: {name: reverse-words@1.0.0}`), and in Python `rl.grpo_lora(environment="reverse-words@1.0.0", ...)`. The whole example, with a GRPO TrainingJob, is in [`examples/training/custom-reward`](https://github.com/nodus-compute/nodus-platform/tree/main/examples/training/custom-reward). Nodus pulls the image from a public registry or from your organization’s space in the Nodus registry; other private registries are not supported for Environments yet.

Any image works if it provides the two commands Nodus runs, as uid 10001 with no network:

* `nodus-env tasks --split train|test --seed N` writes one JSON line per task: `{"taskId", "prompt", "metadata"}`.
* `nodus-env grade` reads `{"taskId", "completion"}` lines and writes one `{"taskId", "verdict", "reward", "evidence"}` line for each, in order.

The Environment becomes `Ready` once its image is verified, with the declared split sizes in `status.splits`. The first TrainingJob that uses a split and seed runs `nodus-env tasks` in one of its own grading Sandboxes and Nodus keeps that manifest for every later run of your organization, so the tasks never change between runs. A version’s image and graders cannot change: publish a new version instead, so earlier results stay reproducible.


---

# Inference

> Call hosted models through an OpenAI-compatible API, route with Indra, and pay per token.

Source: https://nodus-platform-site.pages.dev/docs/guides/inference/
Build revision: 211ad9f836655b1c3a2668c4693e442471f28614

**Beta:** this feature may change.

Nodus inference is an OpenAI-compatible API over hosted open models. Point any OpenAI SDK at the Nodus base URL, use a Nodus API key, and pay per token from your prepaid credits.

## Sign in and create a key

Sign in to the console, open **Settings → API keys**, and create a key with the `inference:invoke` scope.

## Call a model

```python
from openai import OpenAI

client = OpenAI(base_url="https://inference.nodus-compute.ai/v1", api_key="YOUR_NODUS_API_KEY")
reply = client.chat.completions.create(
    model="nodus/gpt-oss-120b",
    messages=[{"role": "user", "content": "What is the capital of France?"}],
)
print(reply.choices[0].message.content)
```

Terminal window

```bash
curl https://inference.nodus-compute.ai/v1/chat/completions \
  -H "Authorization: Bearer $NODUS_API_KEY" -H "Content-Type: application/json" \
  -d '{"model": "nodus/indra", "stream": true, "messages": [{"role": "user", "content": "Hello"}]}'
```

Streaming works as in the OpenAI API (`stream: true`). Completed chat streams end with `data: [DONE]`; an interrupted or errored stream does not receive a generated completion marker. Every response carries a `Nodus-Request-Id` header.

## From the CLI

The CLI calls the same API with the credential of your current context:

Terminal window

```console
$ nodus inference models
MODEL                 NAME           CONTEXT   MAX OUTPUT   PREVIEW
nodus/indra           Indra          -         -            false
openai/gpt-oss-120b   GPT-OSS 120B   128K      64K          false
openai/gpt-oss-20b    GPT-OSS 20B    128K      64K          false
$ nodus inference chat --model nodus/gpt-oss-20b --max-tokens 64 "What is the capital of France?"
The capital of France is Paris.
request ireq_01k6d2x7q9fvjt3y8m0c4r5n2e · openai/gpt-oss-20b · 84 input + 21 output tokens · $0.000015
$ nodus inference receipt ireq_01k6d2x7q9fvjt3y8m0c4r5n2e
```

`chat` prints the answer on standard output and the request id, the model that answered, the tokens and the charge on standard error; `-o json` prints the whole completion. `receipt` shows any request’s charge for 30 days.

## Models

`GET /v1/models` lists the models you can call now. Each model has a catalog name such as `openai/gpt-oss-120b` and a `nodus/` alias such as `nodus/gpt-oss-120b`. Models marked preview have provisional prices.

In the console, open **Inference → Models** and select Indra to send a request to `nodus/indra`, or choose a model from the catalog. The response summary’s **Model** field shows the model identifier submitted with that request, including `nodus/indra` when using Indra. The API’s routing metadata and receipt still record which underlying model served the request.

## Responses, Messages and embeddings

Chat models also accept stateless Responses and Anthropic Messages requests, with the same model ids, credits, endpoint limits and receipt headers. Responses supports text, inline images and client function tools; stored responses, background work and hosted tools are refused. Text Completions accepts one string prompt and one choice, without echo, suffix, logprobs or best-of sampling. Messages token counting is not available.

```python
response = client.responses.create(model="nodus/gpt-oss-20b", input="Hello", max_output_tokens=64, store=False)
print(response.output_text)
vectors = client.embeddings.create(model="nodus/bge-m3", input=["first document", "second document"])
print(vectors.data[0].embedding)
```

Embeddings are available when the embedding model appears in `GET /v1/models`. They charge only input tokens. A model’s `operations` field identifies the routes it accepts, including `Embeddings`, and `aliases` lists its alternative model ids. The console labels embedding models as **text → vectors** and provides embedding examples in the model panel and endpoint API tab. A request can carry up to 2,048 inputs, each within the model’s context window. Messages accepts the same chat models at `/v1/messages`, with `max_tokens`, `messages` and optional `stream: true`; it does not expose private models used by managed agents. Streaming client tool arguments remain intact when several calls interleave.

## Audio

Transcribe or translate an audio file up to 25 MiB (flac, mp3, mp4, mpeg, m4a, ogg, wav or webm) with `openai/whisper-large-v3`, and generate speech with `canopylabs/orpheus-v1-english`:

```python
with open("meeting.m4a", "rb") as f:
    text = client.audio.transcriptions.create(model="nodus/whisper-large-v3", file=f)
print(text.text)

speech = client.audio.speech.create(model="canopylabs/orpheus-v1-english", voice="tara",
                                    input="Your job finished.", response_format="wav")
speech.write_to_file("done.wav")
```

Upload the file itself: audio URLs and streamed transcription are not supported. Transcription and translation are billed per audio hour for the file’s length, with a 10-second minimum per request; speech is billed per million input characters, and its audio is limited to 32 MiB.

## Indra: `nodus/indra`

Send `model: "nodus/indra"` and Indra picks the catalog model best suited to each request: small, fast models for simple requests and stronger models for hard ones. The `X-Nodus-Routed-Model` response header names the model that answered. If routing is unavailable, Indra answers with its default model, so requests never fail because of routing. Indra was first called Composer, and `nodus/auto` still works as another name for it.

Indra comes in two tiers, and the `X-Nodus-Composer-Tier` response header (`free` or `paid`) names the one that answered:

* **Free Indra** is for every organization without the Indra plan. It routes among the three cheapest catalog models that are available (by blended price, three input tokens to one output token) and never draws on your credits. It is limited per organization to 100 requests and 200,000 tokens per UTC day, and each request to 4,096 output tokens (fewer when the day has less left); past either daily limit, requests get `429` with code `QuotaExceeded` until 00:00 UTC. Free Indra can also pause for the rest of the day when it is in heavy demand.
* **Paid Indra** comes with the $20 per month Indra plan. It routes across every catalog model, and you pay the routed model’s per-token price plus the small routing charge, from the plan’s monthly Indra allowance first and then from your credits.

Calling a model by name, directly or through a named endpoint, is always billed per token from your credits.

## Named endpoints and limits

An `InferenceEndpoint` gives a model its own base URL and access policy:

Terminal window

```console
$ nodus create inferenceendpoint support-bot --model nodus/gpt-oss-120b --rpm 120 --max-concurrent 8 --max-cost 50
$ nodus wait ep/support-bot --for=jsonpath={.status.phase}=Running
$ nodus get ep
NAME          PHASE     MODEL                READY   RPM   CONCURRENT   MAX-COST   AGE
support-bot   Running   nodus/gpt-oss-120b   True    120   8            $50        12s
$ nodus inference chat --endpoint support-bot "Where is my order?"
```

Call it at `https://inference.nodus-compute.ai/endpoints/<project>/support-bot/v1`, or send `model: "endpoint/support-bot"` on the shared base URL. Endpoints enforce `rpm`, `tpm`, `maxConcurrent`, `allowedKeys` (API key names, `--allowed-key`) and an optional `maxCostUSD`: once the requests through an endpoint have spent its cap, the next one gets `402 BudgetExceeded`, and raising `maxCostUSD` lets requests through again (it can only be raised). A new endpoint answers `503` for the few seconds until it is `Running`. `nodus stop ep/support-bot` makes it answer `503` until `nodus start ep/support-bot`; its `Ready` condition says whether its model can serve now.

## Billing per token

* Before a request runs, Nodus holds its maximum cost: the input bound plus `max_tokens` at the model’s rates. Lower `max_tokens` to hold less.
* When it finishes, you are charged the tokens used (input, cached input and output) and the rest of the hold is released. A stream you stop early is charged for the tokens it produced.
* Send an `Idempotency-Key` header to retry safely: a repeat within 24 hours returns the same answer with `Idempotent-Replayed: true` and no second charge. A request that failed without a charge (`Released`) frees its key, so the retry runs again. Reusing a key for a different request body returns `409`.
* `GET /v1/requests/{id}` returns the receipt for 30 days: its `operation`, `usage`, `amountUSD`, `pricebookVersion` and `state`: `Running`, `Settled`, `Released` (no tokens, no charge) or `Unknown`. A request whose outcome Nodus never learned is `Unknown` and is never charged.

## Errors

Errors use the OpenAI shape, with the Nodus error code in both `type` and `code`:

|Status|Code|Meaning|
|-|-|-|
|400|`Unsupported`|A feature outside per-token billing: built-in or server-side tools, a non-default `service_tier`, `store`, `background`, file parts, image URLs or audio URLs, a key repeated in another letter case, or an operation the model does not serve|
|400|`Invalid`|A malformed request, or an upload that is not audio in a supported format|
|401|`Unauthorized`|Missing or invalid API key|
|402|`InsufficientCredits`, `BudgetExceeded`|Not enough credit for the hold, or a budget or endpoint cap blocks it|
|409|`RequestInProgress`|The same `Idempotency-Key` is still running; retry after `Retry-After`|
|409|`IdempotencyKeyReused`|The `Idempotency-Key` was used for a different request; send a new key|
|413|`RequestEntityTooLarge`|A JSON body over 1 MiB or an audio file over 25 MiB|
|429|`TooManyRequests`|An org or endpoint limit; retry after `Retry-After`|
|502, 503|`Unavailable`|The model failed, is at capacity or is not available; retry shortly. You are not charged|


---

# Multi-node training (Beta)

> Run one training job across several machines with torchrun, Ray or your own launcher, and understand what it costs to assemble them.

Source: https://nodus-platform-site.pages.dev/docs/guides/multi-node/
Build revision: 211ad9f836655b1c3a2668c4693e442471f28614

**Beta:** this feature may change.

Beta

Multi-node training is in Beta and on for every organization, with no access request and no purchase needed. Beta gangs have at most 8 nodes. When members on different providers connect over their public addresses (the `public` path), traffic between them, including gradients and the rendezvous, is **not encrypted**; each member accepts it only from the other members’ addresses. Keep `network: Colocated` if your data must not cross the internet unencrypted.

A distributed Job runs your command on several machines at once, called a **gang**. Nodus acquires every member, connects them, checks that they can reach each other, and only then starts your command on all of them together. You write an ordinary training script; the launcher finds its peers from environment variables Nodus sets.

## Run a two-node job

train.py

```python
"""A two-node DDP smoke run. torchrun reads its rendezvous from the PET_* variables Nodus sets on every rank."""
import os

import torch
import torch.distributed as dist
from torch.nn.parallel import DistributedDataParallel as DDP

backend = "nccl" if torch.cuda.is_available() else "gloo"
dist.init_process_group(backend)
local_rank = int(os.environ["LOCAL_RANK"])
device = torch.device("cuda", local_rank) if backend == "nccl" else torch.device("cpu")
if backend == "nccl":
    torch.cuda.set_device(device)

# Every process contributes 1, so the sum is the world size.
one = torch.ones(1, device=device)
dist.all_reduce(one)

model = DDP(torch.nn.Linear(16, 1).to(device), device_ids=[local_rank] if backend == "nccl" else None)
opt = torch.optim.SGD(model.parameters(), lr=0.1)
for step in range(20):
    x = torch.randn(32, 16, device=device)
    loss = (model(x) - x.sum(dim=1, keepdim=True)).pow(2).mean()
    opt.zero_grad()
    loss.backward()
    opt.step()

if dist.get_rank() == 0:
    print(f"world_size={int(one.item())} node={os.environ['NODUS_NODE_RANK']} "
          f"transport={os.environ['NODUS_GANG_TRANSPORT']} loss={loss.item():.4f}")
dist.destroy_process_group()
```

run.sh

```sh
nodus run --name torchrun-2-node --gpu H100 --nodes 2 --launcher torchrun --max-cost 2.00 -- torchrun train.py
```

`torchrun` needs no flags: its node count, rendezvous address and local address come from the `PET_*` variables on every rank. The run prints `world_size=2` from rank 0 when both nodes joined.

## Topology

Say how big the gang is with exactly one of:

|Field (`nodus run` flag)|Meaning|
|-|-|
|`distributed.nodes` (`--nodes`)|The number of machines, 2 to 8|
|`distributed.totalGPUs` (`--total-gpus`)|The total GPU count; Nodus picks the machine shape. A total one machine can hold becomes a single-node Job|
|`distributed.gpusPerNode` (`--gpus-per-node`)|GPUs per machine: 1, 2, 4 or 8. With `nodes` it defaults to the GPU count of `resources.gpu`; with `totalGPUs` Nodus picks it unless you set it|

Every member gets the same GPU count. When `resources.gpu.type` lists several accelerator types, members may get different ones, each billed at its own rate. The resolved shape is written once to `status.topology` and never changes on a restart, so the world size stays the same for the whole run. A dry run shows it, with the hold for every member (`gangHoldUSD`) and the assembly bound:

Terminal window

```console
$ nodus run --dry-run --gpu H100 --total-gpus 16 --launcher torchrun -- torchrun train.py
```

`distributed.network` bounds where members may be:

|`network`|Members are placed|
|-|-|
|`Colocated` (default)|With one provider in one region, on its private network|
|`Regional`|With any providers inside one region class|
|`Global`|Anywhere. The run shows the `WANBound` condition, because training between distant machines is limited by the network|

`distributed.transport` says which network paths you accept. `Direct` (default) admits every path that is not relayed: a provider’s private network, the members’ public addresses between providers (unencrypted in the Beta), and direct encrypted paths. `Auto` also admits **relayed** paths, which carry traffic through Nodus relays: they work between providers that cannot reach each other directly, but they are slow and fit small models only. The estimate warns `RelayedLowBandwidth` when a relayed path is possible.

run.sh (across providers)

```sh
nodus run --name torchrun-2-node-xp --gpu H100 --nodes 2 --launcher torchrun \
  --network global --transport auto --startup-timeout 15m --max-cost 2.00 -- torchrun train.py
```

## Launchers

|`launcher`|What each rank runs|
|-|-|
|`Plain` (default)|Your command once per node. `RANK`, `LOCAL_RANK=0`, `LOCAL_WORLD_SIZE=1` and `WORLD_SIZE` (the node count) make `env://` initialization work for one process per node|
|`Torchrun`|Your `torchrun` command on every node; rank 0 hosts the rendezvous. Nodus owns restarts, so `PET_MAX_RESTARTS=0`|
|`Ray`|A Ray cluster: rank 0 is the head, the others join it, and your command runs on rank 0 once all nodes are up, with `RAY_ADDRESS` set|
|`Verl`|The Ray launcher plus `NODUS_VERL_OVERRIDES` (`trainer.nnodes`, `trainer.n_gpus_per_node`) to append to your verl command|
|`Accelerate`, `Deepspeed`|The environment contract plus `ACCELERATE_*` or `/etc/nodus/hostfile`. Accepted in the Beta, qualified later|

Every rank gets the same variables except the rank-specific ones:

|Variable|Value|
|-|-|
|`NODUS_NODE_RANK`, `NODE_RANK`|The Nodus rank of this node; rank 0 hosts the rendezvous|
|`NODUS_NUM_NODES`, `NNODES`, `NODUS_GPUS_PER_NODE`|The node count and GPUs per node|
|`NODUS_NODE_IPS`|Every member’s address in rank order|
|`MASTER_ADDR`, `MASTER_PORT`|Rank 0’s address and `29500`|
|`NODUS_GANG_EPOCH`, `NODUS_GANG_TRANSPORT`|The current epoch and its path: `private`, `public`, `direct` or `relayed`|
|`NODUS_RESTORE_URI`|The checkpoint to resume from after a restart, when one exists|

`NODUS_NODE_RANK` is Nodus’s rank. torch assigns its own global `RANK` inside torchrun and may order nodes differently. Setting any `NODUS_*`, `PET_*` or `MASTER_*` variable, `NODE_RANK` or `NNODES` in your spec is rejected. Your `NCCL_*` settings are kept, except the few the network path decides.

## Failures and restarts

If a member’s machine is lost or reclaimed, Nodus stops the whole gang’s current epoch at once and restarts it: surviving machines are kept, only the lost ranks get new machines, and every rank starts again at the next epoch from the latest gang checkpoint (`NODUS_RESTORE_URI`). A process left over from the old epoch cannot join the new rendezvous. `recovery.maxAttempts` counts these restarts (default 3). The first failure’s cause is the gang’s reason, so a crash on one rank that brings down the others reports that crash.

If a machine is refused or never appears while the gang is still assembling, only that rank gets another machine (up to two per assembly) and the registered ranks keep waiting at the start barrier. A suspend or cancel in progress is never turned into a restart: a machine lost during it simply completes the stop.

If your command exits non-zero on any rank, the run fails; a zero exit on every rank succeeds it.

Each rank’s attempts carry the `nodus.dev/rank` label, so `nodus get attempts -l nodus.dev/job=<name>,nodus.dev/rank=1` lists rank 1’s attempts across restarts, with each attempt’s placement, boot or restore time and cost.

## The assembly bound

Assembly waits at most `distributed.startupTimeout` (default 15 minutes, 5 to 60) for every member to become ready and pass the network check. If it does not, every acquired member is released, and Nodus tries again up to `maxAssemblyRetries` times (default 2) before the run fails with `GangAssemblyTimeout`.

You pay for each member from the moment its provider starts billing until it is confirmed deleted, including time spent waiting for the other members. The estimate shows the most a failed assembly can cost, the **assembly bound**: the sum of the members’ rates plus up to two replacement machines, times the startup timeout plus teardown, times the number of tries. Nodus pays, not you, when assembly fails for a Nodus reason: a network check that fails on a path Nodus lists as qualified, an address collision, or an outage of the Nodus mesh.

## Beta caveats

* Public paths carry traffic between members’ public addresses in clear text. Only the gang’s own members may connect to a member’s rendezvous and collective ports, but the bytes are not encrypted on the way.
* Relayed paths carry every byte through a Nodus relay, are billed per relayed GiB, and suit small models only.
* Relayed gangs have at most 2 members until larger relayed gangs are qualified.
* On relayed paths your image must use dynamically linked glibc programs; otherwise the run fails with `ShimNotLoaded`.
* Provider pairs are added as each is qualified; a pair that is not qualified is never offered. The measured TCP throughput and round-trip time per path class are published here as pairs are qualified.


---

# Training (Beta)

> Fine-tune, post-train, distill and pretrain models with TrainingJobs on managed runtimes, on one GPU or across several machines.

Source: https://nodus-platform-site.pages.dev/docs/guides/training/
Build revision: 211ad9f836655b1c3a2668c4693e442471f28614

**Beta:** this feature may change.

Beta

TrainingJob and TrainingRuntime are `nodus.dev/v1beta1`. Fields may still change before they reach v1; a change is announced in the changelog with a migration note. Environments, used for RL and evaluation, are GA.

A TrainingJob says what to train (a base model, a dataset or an environment, and the runtime’s parameters) and Nodus runs it as a Job: it caches the model, picks GPUs from the runtime’s presets, estimates the run before it starts, checkpoints it, and collects the trained weights as outputs.

## Sign in

Terminal window

```console
$ pip install nodus-compute
$ nodus login
```

## Choose a runtime

A TrainingRuntime is a pinned trainer image with a parameter schema, presets and measured step times. The managed runtimes live in the `nodus` catalog project:

|Runtime|What it trains|Data|
|-|-|-|
|`nodus/sft`|Supervised fine-tuning, full or LoRA/QLoRA; `task: Pretrain` for continued or from-scratch pretraining|prompt/completion or text JSONL|
|`nodus/dpo`|Direct preference optimization|prompt, chosen, rejected|
|`nodus/orpo`|Odds-ratio preference optimization (no reference model)|prompt, chosen, rejected|
|`nodus/kto`|Kahneman-Tversky optimization from thumbs-up/down labels|prompt, completion, label|
|`nodus/reward-model`|A reward model for RLHF|prompt, chosen, rejected|
|`nodus/grpo-lora`|GRPO reinforcement learning against an Environment (single node)|an Environment|
|`nodus/distill`|Knowledge distillation from a teacher model|prompts or chats|
|`nodus/evaluate`|Evaluation only (`mode: Evaluate`)|an Environment or a benchmark|

Terminal window

```console
$ nodus get trainingruntimes -n nodus
$ nodus get trainingruntime/sft -n nodus -o yaml   # the parameter schema, presets and step times
```

In the console, **Environments** lists the same runtimes under **Training templates**. You start a run from the SDK or the CLI; the console shows it under **Runs › Post-training** with its stages, curves, results, cost and outputs.

## Fine-tune with LoRA

sft.yaml

```yaml
apiVersion: nodus.dev/v1beta1
kind: TrainingJob
metadata:
  name: support-sft
spec:
  runtime: nodus/sft
  model:
    uri: hf://Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca
    secret: hf-token            # only for gated repositories
  data:
    volume: support-chats       # a Volume holding chats.jsonl
    path: chats.jsonl
    format: JSONL
    columns: {prompt: question, completion: answer}
  parameters:
    steps: 2000
    lora: {r: 16, alpha: 32}
  maxCostUSD: "20.00"
```

Review it before anything runs. A dry run compiles the TrainingJob into the Job it will create and prices it:

Terminal window

```console
$ nodus apply -f sft.yaml --dry-run -o yaml     # status.estimate: cost to completion, first hold, start time
$ nodus apply -f sft.yaml
```

What the dry run decides:

* **Resources** come from the first runtime preset that matches the model, the method (Full, LoRA or QLoRA, read from `parameters`) and quantization. Anything you set under `spec.resources` wins.
* **The estimate** is the number of steps times the measured seconds per step on the slowest GPU type you allowed, plus evaluation tasks times seconds per task. Set `spec.expectedDuration` to use your own figure instead.
* **Parameters** are merged over the runtime’s defaults and checked against its schema. A typo such as `learning_rate` is rejected with the field path, before you are charged anything.

Model, teacher, and dataset Volume inputs are fixed to their latest committed revision when you create the TrainingJob, unless you select an explicit revision. Upload the files and wait for the Volume to be Ready first. Later uploads do not change the admitted run.

Pin Hugging Face model revisions to a 40-character commit. The first TrainingJob that names a revision imports it once into a read-only Volume in your project, and every later run on that revision reuses it. The trainer runs with `HF_HUB_OFFLINE=1`, so billed GPU time is never spent downloading weights.

## Watch it

Terminal window

```console
$ nodus get tj -w
NAME          RUNTIME         PHASE     STAGE        CHANGE   COST    AGE
support-sft   nodus/sft   Running   Training              $0.12   6m
$ nodus describe tj/support-sft
$ nodus logs job/support-sft -f
```

`STAGE` is `CachingModel` while the model imports, then the stage the trainer reports (`Baseline`, `Training`, `Evaluating`), and `Exporting` while the outputs are written to `spec.exportTo`. `describe` adds the step, total steps, loss and tokens per second, the provenance of the run (runtime and environment digests, model revision, and whether the runtime is a managed one) and the cost so far.

In the console, open the run from **Compute**: the Overview shows the stages, step, loss and the baseline and final pass rates; **Gradings** lists every graded task as it is scored; **Results** has the outputs and the before and after table.

The TrainingJob creates a Job of the same name and owns it. Logs, checkpoints and outputs are the Job’s, so everything in the [Jobs guide](https://nodus-platform-site.pages.dev/docs/guides/jobs/) applies.

## Download the results

Set `spec.exportTo.volume` to a Ready `ReadWriteOnce` Volume in the same project to also copy the complete result there. The runtime must declare an output at `/nodus/outputs`. The export uses an isolated CPU Job under the TrainingJob’s spending cap, waits up to 5 minutes for capacity, and allows 15 minutes for the copy. It follows suspension and cancellation. Success records the exact committed Volume revision in the `Exported` condition; files remain downloadable as Job outputs if the copy fails. Source or destination symbolic links are rejected.

When the run succeeds, every file the trainer wrote under `/nodus/outputs` is an output of the TrainingJob, named by its path: `results.json` (the final metrics), `provenance.json`, `manifest.json` (a sha256 per file) and the weights under `adapter/` (LoRA and QLoRA) or `model/` (full fine-tuning), with the tokenizer beside them.

Terminal window

```console
$ nodus get tj/support-sft -o jsonpath='{.status.outputs[*].name}'
$ nodus cp tj/support-sft:outputs/results.json ./results.json
$ nodus cp tj/support-sft:outputs/adapter/adapter_model.safetensors ./adapter_model.safetensors
```

`nodus cp` verifies each file against the sha256 recorded when it was collected.

## Full fine-tuning and pretraining

Leave out the `lora` block for full fine-tuning; presets for the Full method pick larger GPUs.

Pretraining uses the same runtime with `task: Pretrain`. `initialization: Continued` (the default) continues from the model’s weights; `Scratch` uses only its architecture and tokenizer:

```yaml
spec:
  runtime: nodus/sft
  task: Pretrain
  model: {uri: hf://Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca}
  data: {volume: corpus, format: JSONL}
  parameters: {steps: 10000, initialization: Scratch, packing: true}
  resources: {gpu: {type: [h100-sxm-80g], count: 8}}
```

## Preference tuning

`dpo`, `orpo`, `kto` and `reward-model` take a dataset whose columns you map onto the names the runtime expects:

```yaml
spec:
  runtime: nodus/dpo
  model: {volume: {name: my-sft-model, revision: 3}}   # a model you trained earlier, by Volume
  data:
    volume: prefs
    format: JSONL
    columns: {prompt: prompt, chosen: chosen, rejected: rejected}
  parameters: {steps: 400}
```

A model can come from a Hugging Face URI or from one of your Volumes at a pinned revision, so a DPO run can start from the output of an SFT run.

## Distillation

`distill` trains the student in `model` on the teacher’s distributions. The teacher is cached the same way:

```yaml
spec:
  runtime: nodus/distill
  model: {uri: hf://Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca}
  teacher: {uri: hf://Qwen/Qwen3-4B@0e9e39f249a16976918f6564b8830bc894c89659}
  data: {volume: chats}
  parameters: {temperature: 2.0}
  resources: {gpu: {type: [h100-sxm-80g], count: 2}}
```

## Reinforcement learning (Beta, single node)

`grpo-lora` trains against an [Environment](https://nodus-platform-site.pages.dev/docs/guides/environments/): it samples completions for the environment’s tasks and Nodus grades them. The trainer never sees the expected answers and never grades itself: it submits completions to the TrainingJob’s `gradings` endpoint, and Nodus runs the environment’s grader in isolated containers on the same GPU host, with no network access, model files or GPU devices and the answers readable only by the grader. Task generation also uses this host; Nodus does not provision a separate CPU machine. A grader that fails is reported as an infrastructure failure and never scores the completion 0.

Environment-backed training and evaluation require GPU capacity that supports isolated grading. Grading is serialized per TrainingJob. Baseline and final model evaluation run on the GPU. Custom runtime images must implement the task-fetch endpoint and declare `spec.environmentTasks: AttemptAPI`; older runtimes are rejected before GPU acquisition.

```yaml
spec:
  runtime: nodus/grpo-lora
  model: {uri: hf://Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca}
  environment: {name: nodus/graph-coloring@1.0.0, trainTasks: 50, heldOutTasks: 64, seed: 42}
  evaluation: {baseline: true, final: true}
  parameters: {steps: 50, lora: {r: 8}}
  resources: {gpu: {type: [rtx4090-24g], count: 1}}
  maxCostUSD: "5.00"
```

With `evaluation.baseline` and `evaluation.final`, the run scores the held-out tasks before and after training. `status.summary` holds both pass rates and a comparison: the change in percentage points, and whether it is comparable at all. A comparison needs the same task set on both sides, at least 16 scored tasks, and a result that is not all-pass or all-fail on both sides; otherwise its `reason` says why (`TooFewTasks`, `DifferentTasks`, `NonePassed`, `AllPassed`, `IncompleteEvents`). `status.diagnostics.rewardSignal` warns when every training reward was the same, which teaches a group-relative method nothing.

### Reproducing the legacy letter-counting protocol

Use `nodus/letter-counting-legacy-eval@1.0.0` for the legacy evaluation protocol: the first 64 tasks from Reasoning Gym’s `letter_counting` generator with seed 42, an answer-only system message, greedy decoding with 256 completion tokens, and the generator’s raw-completion scorer. The `letter-counting` catalog example pins Qwen3-1.7B and generates one prompt at a time.

Use `nodus/letter-counting-legacy-rl@1.0.0` for the legacy RL protocol. Its first 64 tasks are training data; the following 64 are held out for both baseline and final evaluation. The example sets 50 optimizer steps, learning rate `5e-6`, eight generations, generation batch size eight, per-GPU batch size one, gradient accumulation four, temperature `1.2`, top-p `0.95`, top-k `50`, and 64 completion tokens. It uses a rank-eight LoRA adapter with alpha 16, no dropout, all linear layers, bf16, non-reentrant gradient checkpointing, and a linear learning-rate schedule without warmup. Thinking is disabled in both phases. Recovery checkpoints remain enabled.

These profiles pin Reasoning Gym `0.1.25` and preserve the legacy source’s task slicing, prompt roles and scoring. They do not establish historical binary equivalence or a measured pass rate. The evaluation profile’s 64 tasks are different from the RL profile’s held-out 64 tasks; compare baseline and final within the same RL run. The existing `nodus/reasoning-gym` profile uses shuffled family-specific seeds and remains a different benchmark. Record the exact runtime, environment and model digests alongside results before claiming a numeric comparison.

RL runs on one node. Multi-node RL and custom environment collections are not available yet.

## Multi-node training

Set `resources.nodes` above 1 to train across several machines, each with `resources.gpu.count` GPUs. The TrainingJob compiles into one multi-node Job (a gang): every node starts together, `torchrun` is launched on each with its rank and the rendezvous address, and the run checkpoints with `torch.distributed.checkpoint`.

```yaml
spec:
  runtime: nodus/dpo
  model: {volume: {name: my-8b, revision: 3}}
  data: {volume: prefs, format: JSONL, columns: {prompt: prompt, chosen: chosen, rejected: rejected}}
  parameters: {steps: 400}
  resources: {gpu: {type: [h100-sxm-80g], count: 8}, nodes: 2}
  distributed: {network: Global, transport: Auto}
```

* **Launchers.** A runtime declares `Torchrun`, `Accelerate`, `Deepspeed` or `Plain`, and the run uses that launcher on every node: `torchrun` reads the gang’s rendezvous settings, `accelerate launch` and DeepSpeed’s per-node launcher read the rank, address and world size Nodus sets on each node, and `Plain` starts the runtime’s command once per node. A single node runs `torchrun --standalone` for a `Torchrun` runtime, `accelerate launch` or `deepspeed` for the other two.
* **Placement.** `distributed.network` is `Colocated` (one provider and region, the fastest interconnect), `Regional` or `Global`; `transport: Auto` also allows an encrypted SSH connection between two restricted GPU machines without a separate relay server. Native GPU-to-GPU connectivity is preferred. Wider settings find capacity sooner at some cost in throughput.
* **Recovery.** If a node is lost, the whole gang restarts from the latest checkpoint every rank committed, up to 3 times; progress since that checkpoint is redone. A single-node run restores its latest snapshot instead. Managed trainers bind recovery state to the full run configuration, runtime image digest, and pinned model and dataset references; incompatible state is refused. GRPO reuses the baseline saved in that state so recovery does not repeat baseline grading.
* **Cost.** The estimate covers every node. Assembly time (nodes waiting for each other) is billed and bounded; see [Multi-node training](https://nodus-platform-site.pages.dev/docs/guides/multi-node/) for the limits.

## Spending caps

`spec.maxCostUSD` caps the whole TrainingJob: its Job, its same-host grading, its output export and its retries. When the cap is reached the run is suspended, not failed, and active grader containers are removed. Raise the cap to resume the same run from its checkpoint:

Terminal window

```console
$ nodus patch tj/support-sft --type merge -p '{"spec":{"maxCostUSD":"40"}}'
```

The cap can only be raised. `spec.state: Suspended` pauses a run yourself, and `Running` resumes it.

## Tracking

`spec.tracking.mlflow` (`uri`, `secret`, `experiment`) or `spec.tracking.wandb` (`connection`) send the trainer’s metrics to your own MLflow server or Weights & Biases project as well. Nodus shows its own stage and metrics either way. The catalog runtimes report to MLflow when the run has `MLFLOW_TRACKING_URI` and to W\&B when it has `WANDB_API_KEY`, which is what those two fields set; the Secret holds `MLFLOW_TRACKING_TOKEN` (or `MLFLOW_TRACKING_USERNAME` and `MLFLOW_TRACKING_PASSWORD`), and the W\&B Connection’s Secret holds `WANDB_API_KEY`.

## Bring your own runtime

A TrainingRuntime in your project can run any image. The trainer reads:

|Variable|Contents|
|-|-|
|`NODUS_INPUT_PARAMS`|a directory that holds `params.json`, the run: `task` and `mode` (for example `SFT`, `Train`), the merged, schema-checked `parameters`, `data` (path, format and column mapping), `environment` (task counts, seed and graders), `evaluation`, `trainingJob` and the pinned `modelRevision` and `teacherRevision`|
|Environment tasks|Call `POST …/trainingjobs/{name}/tasks` through the attempt API socket. The response `files` maps `train.jsonl` and `test.jsonl` to JSONL strings containing `{taskId, prompt, metadata}` and no hidden answers. Managed runtimes fetch them automatically; custom runtimes must implement this step.|
|`NODUS_INPUT_MODEL`, `NODUS_INPUT_TEACHER`, `NODUS_INPUT_DATA`|where the model, teacher and data are mounted|
|`NODUS_STATE_DIR`, `NODUS_OUTPUT_DIR`|resumable state and results|
|`NODUS_NODE_RANK`, `MASTER_ADDR`, `PET_*`|the gang of a multi-node run|
|`NODUS_CHECKPOINT_URI`, `NODUS_RESTORE_URI`|where ranks write and restore distributed checkpoints|

Write and load your own model, optimizer and progress files under the state directory; Nodus decides when to checkpoint and where checkpoints are stored, and never adds resume flags to your command.
