This is the full developer documentation for Nodus
# Build with Nodus
> Run your first job, choose a compute feature, and get from code to results.
Run your code on cloud CPUs and GPUs. Start with a command, a Python function, or an interactive environment.
> **From code to your first result**
>
> Install the CLI, sign in, and run a job in about five minutes.
>
> [Run your first job →](/docs/getting-started/)
## What are you building?
[Section titled “What are you building?”](#what-are-you-building)
Choose a starting point. Each guide covers the essentials and a working example.
* **[Jobs](/docs/guides/jobs/)** Run a script or batch command to completion.
* **[Workspaces](/docs/guides/workspaces/)** Develop with SSH, VS Code, or Jupyter.
* **[Functions](/docs/guides/functions/)** Call Python remotely. Run calls in parallel.
* **[Sandboxes](/docs/guides/sandboxes/)** Give agent-written code an isolated place to run.
* **[Inference](/docs/guides/inference/)** Call a hosted model from your application.
* **[Agents](/docs/guides/agents/)** Run agents with tools and saved progress.
Looking for [training (Beta)](/docs/guides/training/), [storage](/docs/guides/volumes/), or [billing](/docs/guides/billing/)? [Browse all guides →](/docs/guides/)
## Bring your coding agent
[Section titled “Bring your coding agent”](#bring-your-coding-agent)
[Connect over MCP](/docs/for-agents/) to work from your editor, or give your agent the [Markdown task map](/docs/source/index.md). Read one relevant guide at a time; use the [reference](/docs/reference/) for exact commands and API fields.
# For coding agents
> Connect a coding agent over MCP, and the machine-readable versions of these docs, the API and the setup.
Coding agents use Nodus the way you do: the same account, projects, budgets and confirmations. Connect one over MCP, or point it at the machine-readable outputs below.
## Read in this order
[Section titled “Read in this order”](#read-in-this-order)
1. Read the [compute decision guide](/docs/source/guides/choose-compute.md) to choose the feature that fits the request.
2. If this is a first run, read the [quickstart](/docs/source/getting-started.md).
3. Use the [guide directory](/docs/source/guides.md) to find the relevant guide, then fetch its Markdown.
4. Look up exact commands in the [CLI reference](/docs/reference/cli/), signatures in the [Python reference](/docs/reference/python/), or request fields in [OpenAPI](/docs/openapi.json).
Fetch individual pages to keep context small. For example, `/docs/guides/jobs/` has its Markdown at [`/docs/source/guides/jobs.md`](/docs/source/guides/jobs.md). Each docs page has a **View Markdown** link. Use the full corpus only for tasks that need many parts of the product.
## Connect over MCP
[Section titled “Connect over MCP”](#connect-over-mcp)
Nodus runs one MCP server with generic tools over every resource (`get`, `describe`, `logs`, `estimate`, `apply`, `exec` and more). Writes return a dry run first and run only once confirmed.
| Client type | Configuration |
| ---------------------------------------------------------- | -------------------------------------------------------------------------------------------- |
| Hosted (Claude Code, Codex, Cursor and other HTTP clients) | Add the server URL from [`/mcp-hosted.json`](/mcp-hosted.json) and sign in with your browser |
| Local stdio clients | Run `nodus mcp`; the configuration is [`/mcp.json`](/mcp.json) |
Step-by-step setup for each client is at [/connect](/connect/), and as Markdown at [`/connect.md`](/connect.md).
## Machine-readable outputs
[Section titled “Machine-readable outputs”](#machine-readable-outputs)
| Output | What it holds |
| ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------- |
| [`/docs/pages.json`](/docs/pages.json) | Lightweight page manifest: titles, summaries, status and source URLs |
| [`/docs/llms.txt`](/docs/llms.txt) | An index of these docs for language models |
| [`Start`](/docs/bundles/start.md), [`compute`](/docs/bundles/compute.md), [`models`](/docs/bundles/models.md), [`data`](/docs/bundles/data.md), [`billing`](/docs/bundles/billing.md) | Focused bundles with complete pages, examples and source URLs |
| [`/llms-full.txt`](/llms-full.txt) | Guides and concepts in one Markdown file; use the index for exact references |
| `/docs/source/.md` | Each page’s Markdown, linked from the page with `rel="alternate"` |
| [`/docs/index.json`](/docs/index.json) | Every page with its headings and text, versioned by build |
| [`/docs/openapi.json`](/docs/openapi.json) | The HTTP API contract |
| `/skills//SKILL.md` | Task instructions for agents that support skills |
| [`/install`](/install), [`/install.ps1`](/install.ps1) | The CLI installer for macOS, Linux and Windows |
A good first prompt for an agent that can read URLs:
```text
Read https://nodus-compute.ai/connect.md and help me connect Nodus to this agent. Reuse any existing Nodus
connection. Verify setup by listing my jobs. Do not start paid compute.
```
## Use the current contract
[Section titled “Use the current contract”](#use-the-current-contract)
Fetch the relevant guide and its linked reference before writing commands. Keep the resource’s API version and Beta status explicit. Use the published OpenAPI schemas for field names and the CLI or SDK reference for the installed interface; do not invent flags from examples for another tool.
Each Markdown page includes its canonical source URL and build revision. Cite the source page when explaining behavior. Read billing and recovery limits before creating work, and verify the resource status, logs and outputs before reporting success. A submitted request is not evidence that the run completed.
# Get started
> Install the Nodus CLI, sign in, run a command on a GPU, follow it, get its results and see exactly what it cost.
This page takes you from nothing to a finished GPU job and its bill. You need a terminal and a browser. New accounts start with a **$30 starter grant**, so the first runs need no card.
1. **Install the CLI.**
* pip
```sh
pip install nodus-compute
```
The Python package (Python 3.10 or newer) includes the `nodus` CLI.
* macOS and Linux
```sh
curl -fsSL https://nodus-compute.ai/install | sh
```
* Windows
```powershell
irm https://nodus-compute.ai/install.ps1 | iex
```
2. **Sign in.**
```sh
nodus login
```
Your browser opens to confirm the sign-in. A new account gets an org with a `default` project, an API key for this machine, and the $30 starter grant, which expires 30 days after it is granted. On a machine without a browser, run `nodus login --device` and approve the code from any other device.
3. **Run a command on a GPU.**
```console
$ nodus run --gpu L4 --image nodus/pytorch -- python -c "import torch; print(torch.cuda.get_device_name())"
job/run-4kq7z created · est. $0.01–0.03 · starts in ~2–4 min (cold) · Ctrl+C to cancel, -d to detach
✓ Scheduled l4-24g-x1-us · $0.52/h (rate frozen)
✓ Provisioning 1m48s
✓ Pulling image 22s
▶ Running
NVIDIA L4
✓ Succeeded in 2m31s · $0.02 · kept warm 60 s · nodus describe job/run-4kq7z
```
`nodus run` uploads the current directory, prints the estimate before anything is charged, streams the logs and exits with your command’s exit code. The rate is fixed when the machine is chosen and holds for the whole run. Add `-d` to return right away and let the job run on its own.
4. **Follow it.** Every run is a Job you can come back to.
```sh
nodus get jobs -w # live list of your jobs
nodus logs -f job/run-4kq7z # stream the logs again
nodus describe job/run-4kq7z # status, events and the cost so far
```
5. **Get the results.** Declare an output path when you run, then copy it back. Downloads are checked against their SHA-256 digest.
```sh
nodus run --gpu L4 --image nodus/pytorch --output model=/outputs/model -- python train.py
nodus cp job/:outputs/model ./model
```
6. **See what it cost.**
```console
$ nodus billing
$ nodus get usage --field-selector object.name=run-4kq7z --group-by segment
SEGMENT AMOUNT
Boot $0.018778
Running $0.003033
Teardown $0.002889
```
You pay for the machine from the moment it is created until it is deleted: starting up (`Boot`), your command (`Running`) and shutting down (`Teardown`), all at the rate frozen at launch. Credit is prepaid: a hold is reserved before a machine starts, and only what was used is charged.
Using a coding agent?
Connect Claude Code, Codex or Cursor to Nodus at [/connect](/connect/), then ask it to run your command. The agent uses the same account, projects and spending limits as the CLI.
## Next steps
[Section titled “Next steps”](#next-steps)
* [Concepts](/docs/concepts/): resources, projects, and how billing is measured.
* [Pricing reference](/docs/reference/pricing/): list prices for every GPU, CPU shape, storage and egress.
* [Guides](/docs/guides/): one guide per feature.
# Find a guide
> Choose a task, then read the guide you need. Start with running code and add storage, integrations, or billing as needed.
New to Nodus? Start with [your first job](/docs/getting-started/). Otherwise, choose the task you have now.
Not sure which feature fits? [Compare Jobs, Workspaces, Sandboxes and Functions](/docs/guides/choose-compute/).
## Run code
[Section titled “Run code”](#run-code)
| Task | Guide |
| --------------------------------------- | --------------------------------------- |
| Run a command to completion | [Jobs](/docs/guides/jobs/) |
| Work in SSH, VS Code, or Jupyter | [Workspaces](/docs/guides/workspaces/) |
| Run code in an isolated container | [Sandboxes](/docs/guides/sandboxes/) |
| Call Python remotely or in parallel | [Functions](/docs/guides/functions/) |
| Chain jobs together | [Pipelines](/docs/guides/pipelines/) |
| Try multiple parameter combinations | [Sweeps](/docs/guides/sweeps/) |
| Choose resources and check availability | [Compute selection](/docs/guides/gpus/) |
## Models and agents
[Section titled “Models and agents”](#models-and-agents)
| Task | Guide |
| ---------------------------------------- | ------------------------------------------------------ |
| Call a hosted model | [Inference](/docs/guides/inference/) |
| Run agents with tools and saved progress | [Agents](/docs/guides/agents/) |
| Fine-tune or train a model | [Training (Beta)](/docs/guides/training/) |
| Train across machines | [Multi-node training (Beta)](/docs/guides/multi-node/) |
| Define training and evaluation tasks | [Environments](/docs/guides/environments/) |
## Files and state
[Section titled “Files and state”](#files-and-state)
| Task | Guide |
| -------------------------------------------- | ---------------------------------------- |
| Package code and dependencies | [Images](/docs/guides/images/) |
| Keep files between runs | [Volumes](/docs/guides/volumes/) |
| Pass credentials to your code | [Secrets](/docs/guides/secrets/) |
| Download results | [Outputs](/docs/guides/outputs/) |
| Save and load progress after an interruption | [Checkpoints](/docs/guides/checkpoints/) |
| Connect external data and services | [Connections](/docs/guides/connections/) |
## Tools and integrations
[Section titled “Tools and integrations”](#tools-and-integrations)
* [Python SDK](/docs/guides/python/): write and manage runs from Python.
* [MCP](/docs/guides/mcp/): connect your coding agent to Nodus.
* [Console](/docs/guides/console/) and [Ask Nodus](/docs/guides/assistant/): manage work in the browser.
* [Logs and metrics](/docs/guides/logs/): follow progress and diagnose a run.
* [Webhooks](/docs/guides/webhooks/) and [notifications](/docs/guides/notifications/): react to changes.
## Account, costs, and your own compute
[Section titled “Account, costs, and your own compute”](#account-costs-and-your-own-compute)
* [Account setup](/docs/guides/access/sign-up-and-orgs/), [projects](/docs/guides/access/projects/), and [members and roles](/docs/guides/access/members-and-roles/): organize your team.
* [API keys](/docs/guides/access/api-keys-and-scopes/): give applications access.
* [Billing](/docs/guides/billing/), [usage and costs](/docs/guides/billing/usage-and-costs/), and [budgets](/docs/guides/billing/budgets/): fund work and control spending.
* [Pools](/docs/guides/pools/) and [cloud accounts](/docs/guides/pools/cloud-accounts/): use your own compute.
## Coming from another tool?
[Section titled “Coming from another tool?”](#coming-from-another-tool)
Start with [Nodus for Modal users](/docs/getting-started/modal-users/), [for kubectl users](/docs/getting-started/kubectl-users/), or [migrating from Nodus 0.x](/docs/guides/migrate-from-0x/).
# How Nodus works
> Resources, orgs and projects, how your work is placed and kept alive, and how prepaid billing measures it.
This page is the map. Each section links to the concept page that goes deeper.
## Everything is a resource
[Section titled “Everything is a resource”](#everything-is-a-resource)
You describe work as a **resource**: a small declarative object with a `kind`, a `metadata.name` and a `spec`, the same shape Kubernetes uses. You create it with the CLI, the Python SDK, the console or the HTTP API, and Nodus reports progress in its `status`. The same object reads the same everywhere, so `nodus get`, `kubectl get` and the console show one truth.
| You want to | Kind | Everyday command |
| --------------------------------------------- | -------------------------------------------------------- | ---------------------------------------- |
| Run a command to completion | `Job` (and `Pipeline`, `Sweep` to chain or fan out Jobs) | `nodus run`, `nodus apply -f job.yaml` |
| Keep an isolated container for untrusted code | `Sandbox` | `nodus create sandbox`, `nodus exec -it` |
| Call Python remotely and fan out | `App`, `Function`, `FunctionCall` | `nodus deploy app.py` |
| Run a durable agent | `Agent`, `AgentRun`, `AgentGroup` | `nodus create agentrun` |
| Serve or call a model | `Model`, `InferenceEndpoint` | `nodus get models -n nodus` |
| Train or evaluate with a recipe | `TrainingJob`, `TrainingRuntime`, `Environment` | `nodus create trainingjob` |
| Develop on a remote machine | `Workspace` | `nodus ssh workspace/` |
| Store data, images and secrets | `Volume`, `Image`, `Secret`, `Connection` | `nodus volume put` |
| Bring your own machines | `Pool`, `Node`, `EnrollmentToken`, `CloudAccount` | `nodus create pool` |
`TrainingJob` and `TrainingRuntime` are Beta. Every other kind above is generally available. Run `nodus api-resources` for the full list and `nodus explain .spec` for any field.
## Orgs and projects
[Section titled “Orgs and projects”](#orgs-and-projects)
An **org** is your billing and access boundary: members, API keys, credit and budgets belong to it. Inside an org, **projects** group resources (every org starts with `default`). Pass `-p ` to the CLI, or set `NODUS_PROJECT` for a whole shell. Names are unique within a project, and labels such as `team=nlp` let you select resources and break down cost across projects.
## How your work runs
[Section titled “How your work runs”](#how-your-work-runs)
You state requirements (an accelerator and count, memory, a region class, a deadline or a cost ceiling) and Nodus chooses an **offering** that satisfies them, such as `h100-sxm-80g-x8-us`. You see offerings and Nodus ids, never the machines behind them. Each placement of your container on a machine is an **Attempt**. When capacity is reclaimed, Nodus starts a new Attempt and restores the files your program saved to its checkpoint directory (`NODUS_CHECKPOINT_DIR`). Your program reloads its own model, optimizer and progress from those files; Nodus restores files, not process memory.
## How billing works
[Section titled “How billing works”](#how-billing-works)
Nodus sells **prepaid credit**, and four rules decide every charge:
1. **A hold comes first.** Before any paid machine starts, a hold reserves enough credit on your org, and on every budget that applies, to cover the expected run. A launch that cannot be funded is refused with the amounts and a fix, and a running Job stops gracefully, inside its reserved amount, when money runs out.
2. **The rate is frozen at launch.** `nodus run` and the console show the estimate before launch; the rate chosen when the machine is acquired is the rate for the whole run, and it is never above the published list price.
3. **You pay what the provider bills for your machine, and nothing Nodus caused.** Rented capacity bills from the moment the machine is created until its deletion is confirmed, at provider cost ÷ 0.875 (Nodus keeps 12.5 % of what you pay). Capacity Nodus chose and discarded, orphaned machines and Nodus failures are never charged.
4. **Usage is itemized by segment.** Every compute usage line names the part of the machine’s billed time it covers:
| Segment | Covers |
| ---------- | ---------------------------------------------------------------------------------------- |
| `Boot` | From the machine’s creation to your command starting: start-up, readiness and image pull |
| `Running` | Your command, until it stops |
| `Restore` | A replacement machine’s start-up when your work resumes from a checkpoint |
| `Teardown` | From stop to confirmed deletion, plus the provider’s rounding increment |
`nodus get usage --group-by segment` and the console’s Cost tab show the breakdown for any object.
Nodus-operated capacity (Sandbox nodes, warm pools, CPU nodes) bills from the pricebook’s published per-vCPU, per-GiB and per-disk rates instead. Storage above 10 GB per org and egress above 10 GiB per org per day are metered; logs are free.
The starter grant
The first org a verified user creates receives **$30 of credit that expires 30 days after it is granted**. Grant credit is spent before purchased credit, soonest-expiring first. Until your first purchase, the org has starter limits (one Nodus node, three live Sandboxes, Sandbox lifetimes up to two hours), which lift when you buy credit.
See the [pricing reference](/docs/reference/pricing/) for every published rate.
# Attempts and recovery
> How a run survives a reclaimed or lost machine, what each continuity mode keeps, and who pays for the time a recovery takes.
Every run on Nodus executes as one or more **attempts**. An attempt is one incarnation of your command on one machine. When that machine is reclaimed, loses its network or fails, Nodus starts a new attempt somewhere else and your run continues from what it saved. This page explains what carries over, how Nodus decides a machine is gone, and what you pay for along the way.
## Attempts and epochs
[Section titled “Attempts and epochs”](#attempts-and-epochs)
A Job index, a Sandbox or a worker slot holds **one running attempt at a time**. Each new attempt gets the next **epoch**, a number that only grows. Nodus accepts reports, checkpoints and outputs only from the current epoch, so a machine that comes back after it was replaced can never overwrite the work of its successor.
You see a run’s attempts with `nodus get attempts -l nodus.dev/job=`. A run’s attempts share its name with an epoch suffix, and gang members add a rank suffix (`-r1`, `-r2`).
| Attempt phase | Meaning |
| --------------------------------- | ------------------------------------------------------------------------------------------------- |
| `Pending`, `Placing`, `Acquiring` | Choosing and preparing capacity |
| `Starting` | The machine is ready; the image is pulled and inputs or a checkpoint are restored |
| `Running` | Your command is running |
| `Succeeded` | Your command exited 0 |
| `Failed` | Your command or its machine failed; `reason` says which (`NodeLost`, `Preempted`, `OOMKilled`, …) |
| `Cancelled` | The attempt was stopped on purpose: a suspend, a cancel, a budget or lifetime limit |
## Continuity modes
[Section titled “Continuity modes”](#continuity-modes)
`recovery.continuity` says what a new attempt starts from:
| Mode | A new attempt starts with | Use it for |
| -------------- | --------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------ |
| `Checkpointed` | The latest committed checkpoint of your declared state paths (`NODUS_CHECKPOINT_DIR` by default) | Training and long jobs that save their own model, optimizer and progress files |
| `Restartable` | A cold start plus the progress cursor you reported (`NODUS_CURSOR_COMPLETED`, `NODUS_CURSOR_TOTAL`) | Batch work that can skip what it already finished |
| `Ephemeral` | A cold start | Short or idempotent work |
| `Snapshotted` | The latest filesystem snapshot | Sandboxes |
Restoring files restores **files only**, never process memory. Your program loads its own checkpoint when it starts; Nodus never adds resume flags to your command.
An empty checkpoint never counts as saved progress and never replaces an earlier useful one. If every checkpoint a run commits is empty, the run shows `Checkpointed=False, reason=NotCheckpointable`, and Nodus plans and prices it as `Ephemeral` from then on.
## What survives a preemption
[Section titled “What survives a preemption”](#what-survives-a-preemption)
When a provider reclaims interruptible capacity or a machine stops answering, Nodus recovers **make-before-break**:
1. On a reclaim notice, Nodus asks your attempt for an **urgent checkpoint** and, at the same time, starts preparing a replacement machine.
2. The replacement is prepared up to the point where it could start, but it does **not** start while the old machine might still be writing.
3. It starts only once the old machine is provably gone: the old attempt acknowledged its stop, the provider confirmed the machine terminated, or the old attempt’s lease ran out.
4. The new attempt restores according to the continuity mode above.
So a `Checkpointed` run loses at most the work since its last committed checkpoint, and a checkpoint committed during the reclaim notice still counts. If the old machine comes back before it is replaced, the replacement is released and your run simply continues; you are not charged for the replacement.
A machine that stops sending heartbeats gets a replacement prepared after 20 seconds. After 60 seconds of silence an attempt on shared capacity is declared lost and replaced. An attempt on a machine dedicated to your organization may keep running up to the edge of its funding while its replacement waits, because such a machine can only lose its own work.
## Recovery limits
[Section titled “Recovery limits”](#recovery-limits)
Recovery stops, and the run fails with a typed reason, when:
| Reason | Rule |
| ----------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `NoProgress` | Two attempts in a row made no progress. Progress means the command started and either ran for `recovery.minProgressDuration` (default 2 minutes) or committed a non-empty checkpoint. A preemption after progress never counts |
| `RestoreFailed` | The same checkpoint failed to restore twice |
| `ImagePullFailed` | The image failed to pull on a second machine (a pull failure is retried once elsewhere) |
| `MaxAttemptsExceeded` | `recovery.maxAttempts` recoveries were used (default 8; 3 for distributed Jobs) |
| `Preempted`, `NodeLost` | `recovery.onInterruption: Fail` was set, so the first interruption ends the run |
Failures caused by your command (a non-zero exit, out of memory, an invalid checkpoint) fail fast and are not retried. At most three recoveries per organization prepare capacity at the same time; the others wait their turn and show an Event.
## Suspending
[Section titled “Suspending”](#suspending)
`nodus suspend job/` stops the run after a final checkpoint and releases its compute; `nodus resume` continues at a new epoch from that checkpoint. If the final checkpoint fails or takes longer than `max(10 minutes, 2 × the shutdown reserve)`, the run **keeps running** with `Suspended=False, reason=SuspendFailed` and a `SuspendFailed` Event, so a suspend never throws work away. Stops caused by money (credits, budgets, the maximum cost) or by a lifetime limit do not wait: they stop at the funded edge with whatever checkpoint exists.
A stop always wins over recovery. If the machine is lost while a suspend, cancel or money stop is in progress, the attempt ends `Cancelled` with the stop’s reason and nothing is restarted or billed again; `nodus resume` starts the next epoch as usual.
## Who pays for recovery
[Section titled “Who pays for recovery”](#who-pays-for-recovery)
You pay for the machines your run uses, at the rate frozen when each was acquired, from the moment the provider starts billing through the confirmed deletion of the machine. The bill splits each machine’s time into segments:
| Segment | Covers |
| ---------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `Boot` | Start-up: boot, image pull, readiness, and for gangs the wait at the start barrier |
| `Restore` | Start-up of an attempt that resumes: a recovery, a resume after a suspend, and for gangs the surviving machines’ wait from the restart to the next epoch’s start |
| `Running` | Your command running |
| `Teardown` | Stop to confirmed deletion, plus the provider’s billing increment |
A recovery therefore costs you the replacement’s `Boot` or `Restore` time and the lost machine’s `Teardown`. A machine the provider refuses, or that never appears, before it is ready is replaced with another one, up to six tries per attempt; after that the attempt fails with `ReadinessFailed` and the run’s retry rule applies. Nodus pays, and never charges you, for:
* a replacement released because the old machine came back;
* extra machines prepared to start faster (hedges) that did not win;
* machines that failed Nodus’s own checks after creation, and Nodus internal failures;
* for distributed Jobs, probe failures on a qualified network path and outages of the Nodus mesh.
`nodus billing usage --group-by segment` and the Cost tab show every segment of every attempt.
# How billing works
> Prepaid credits, holds before every paid action, exact per-second charges, and what happens when money runs out.
Nodus is prepaid. You add credits, every paid action reserves funds before it starts, and charges come from the rate frozen when the capacity was acquired. Work never runs on credit you do not have, and it stops gracefully, with its progress saved, when money runs out.
## Your balance
[Section titled “Your balance”](#your-balance)
Your org has one balance, made of buckets:
* **Purchased credit** from top-ups you pay for by card.
* **Credit grants**: the starter credit, promo codes and credits from Nodus support. Each grant is its own bucket and may expire.
Charges draw from grants first, soonest expiry first, then from grants that never expire, then from purchased credit. That way a grant is used before it expires, and purchased credit, the only kind that can be refunded, lasts longest.
`nodus billing` and **Usage & billing → Overview** show:
| Field | Meaning |
| ------------------ | ----------------------------------------------------------------- |
| Available | What you can spend now: your buckets minus open holds |
| Reserved | Open holds, with the objects that hold them |
| Purchased, credits | What is left in each kind of bucket, and the next grant to expire |
| Arrears | Unpaid charges; new work waits until a top-up settles them |
## Holds and captures
[Section titled “Holds and captures”](#holds-and-captures)
Before a Job, Sandbox, Workspace, Function worker, agent run or build starts, Nodus places a **hold**: enough funds for the first stretch of work plus the cost of stopping it cleanly. While the work runs, Nodus captures what it used every 5 minutes and renews the hold for the next stretch. When the work ends, the final capture charges the exact amount and the rest of the hold is released.
The estimate before launch shows the hold, the expected cost range and the minimum charge. A create that cannot be funded fails at once with the amounts, instead of queuing and failing later:
```text
Error from server (InsufficientCredits): job "train-a" needs a $3.20 hold to start (released when it ends); available $1.10.
fix: nodus billing top-up 20, or lower spec.maxCostUSD
```
## What is charged
[Section titled “What is charged”](#what-is-charged)
Machines Nodus acquires for you are charged per second from the moment the provider starts billing until the machine is confirmed deleted, at the rate frozen when it was acquired. Usage itemizes each machine’s time by segment (`Boot`, `Restore`, `Running`, `Teardown`). [What you pay for](/docs/guides/billing/what-you-pay-for/) lists every kind of time and who pays for it, and [Pricing](/docs/concepts/pricing/) explains how rates are set.
## Limits
[Section titled “Limits”](#limits)
Three limits apply to every hold, and the tightest wins:
* **Your balance.** A hold never exceeds what is available.
* **Budgets.** A `Block` Budget over the org, a project or a label selector stops new holds and renewals in its scope when it is exhausted.
* **Object caps.** `spec.maxCostUSD` on an object bounds that object and everything it owns.
When a limit is reached, running work stops gracefully: Jobs suspend after a checkpoint, Sandboxes and Workspaces stop, and agent runs wait. Each stopped object shows `Funded=False` with the reason, and resumes when funds return. If a charge would still go past the limit, Nodus absorbs the difference.
## Low balance
[Section titled “Low balance”](#low-balance)
You get a low-balance warning by email, webhook and a console banner when your available balance falls below $5.00 (configurable) or below what your open holds need for their next renewal. Turn on [auto-recharge](/docs/guides/billing/#auto-recharge) to top up automatically below a threshold.
## Arrears
[Section titled “Arrears”](#arrears)
A charge that arrives after its hold is gone, such as daily storage, can leave unpaid charges. While arrears are open, new holds and uploads are refused with `402 ArrearsOutstanding`; your next top-up or grant settles them first.
# The resource model
> How every Nodus object is named, stored, changed, watched and deleted, and how that maps to Kubernetes tools.
Everything you run on Nodus is an **object**: a Job, a Sandbox, a Function, a Volume, a Budget. Every object has the same shape and the same verbs, so once you know one kind you know them all, and the CLI, the Python SDK, the console, MCP clients and `kubectl` all work the same way.
## The shape of an object
[Section titled “The shape of an object”](#the-shape-of-an-object)
```yaml
apiVersion: nodus.dev/v1
kind: Job
metadata:
name: train-a # unique per project and kind
namespace: default # the project
labels:
app: trainer
spec: # what you want
image: nodus/pytorch:2.8-cuda12.8
command: [python, train.py]
maxCostUSD: "40.00"
status: # what Nodus observed; written by Nodus only
phase: Running
```
* **`metadata.name`** is a DNS label of at most 63 characters, unique within its project and kind. Set `metadata.generateName` instead to get a unique suffix; such a create needs an `Idempotency-Key` header.
* **`metadata.namespace`** is the project. Every org starts with the project `default`. The read-only project `nodus` holds what Nodus publishes, such as base images, which you reference as `nodus/:`.
* **`metadata.uid`** is a stable id such as `job_01j9…`, and **`metadata.resourceVersion`** changes on every write. Send the `resourceVersion` you read with an update to make it conditional: if someone changed the object in between, the update fails with error code `Conflict` and you read it again.
* **Labels** select objects (`-l app=trainer`). Labels and annotations under the `nodus.dev/` prefix are set by Nodus; the label `nodus.dev/created-by` records who created an object and `nodus.dev/launched-by` from where.
## Desired state
[Section titled “Desired state”](#desired-state)
You change what an object does by changing its `spec`. Stopping, suspending and cancelling are values of `spec.state`, not separate actions: `nodus suspend job/train-a` sets `spec.state: Suspended`, and a later `nodus apply` of the same file keeps it suspended. Actions that change nothing in `spec`, such as restarting a Sandbox’s service, are requests: `nodus request restart sandbox/dev` sets the annotation `nodus.dev/restart-requested-at` to the current time, and Nodus acts once per value it has not handled yet.
Most `spec` fields cannot change after create. An update that changes one fails with error code `FieldImmutable`, and the response lists each field. The fields that can change, such as `spec.state` and a raise of `spec.maxCostUSD`, are listed in each kind’s reference.
`status.phase` uses the same words on every kind, so `--field-selector status.phase=Running` means the same thing everywhere. Run-to-completion kinds move through `Queued`, `Provisioning`, `Running` and end in `Succeeded`, `Failed` or `Cancelled`; long-running kinds move through `Pending`, `Starting`, `Running` and `Stopped`. `status.conditions` explain the details, such as the condition `Ready`.
## Create, review, apply
[Section titled “Create, review, apply”](#create-review-apply)
* **Create by name is safe to retry.** Creating an object that already exists with the same spec returns the existing object. A different spec fails with error code `AlreadyExists` and lists the differences.
* **Review before you pay.** Add `?dryRun=All` (`--dry-run=server` in the CLI) to run every check without creating anything. The response shows the object with its defaults and, for kinds that use compute, `status.estimate`: the expected cost, the first hold and the start time. Send the `ETag` of that response as `If-Match` on the real create to launch exactly what you reviewed: if the spec or the price book changed, the create fails with error code `PreconditionFailed` and you review again. If only the hourly rate moved, the create goes ahead and never runs above the reviewed rate plus 10 % (unless you set `placement.maxRateUSDPerHour`); when nothing fits, the object waits in `Queued` and its status says the price is above the estimate.
* **Retries never double-create.** The CLI and SDK send an `Idempotency-Key` with every create. Retrying with the same key returns the first response with the header `Idempotent-Replayed: true`; reusing a key for a different request fails with error code `IdempotencyKeyReused`.
## Watch instead of polling
[Section titled “Watch instead of polling”](#watch-instead-of-polling)
Every list can be watched: `nodus get jobs -w`, or `?watch=true&resourceVersion=` on the API. A watch delivers every change after that point, in order, and never skips one. Watches send newline-separated events (`ADDED`, `MODIFIED`, `DELETED`) as JSON, or as `application/x-ndjson` when you ask for it; with `allowWatchBookmarks=true` you also get a bookmark every 30 seconds to resume from. Nodus keeps 24 hours of changes: resuming from an older point fails with error code `Expired`, and you list again.
Lists return at most 500 objects by default (`limit`, up to 1,000) and a `continue` token for the next page.
## Deleting
[Section titled “Deleting”](#deleting)
`nodus delete job/train-a` stops the work first and removes the object once cleanup is done: while it waits the object shows a `deletionTimestamp`, and the phase `Cancelling` or `Terminating`. Nodus tracks that cleanup with finalizers such as the finalizer `nodus.dev/billing`, which clears when the final charge has posted, so you never see an object gone while it is still costing money. Deleting an object also deletes what it created (a Pipeline’s Jobs, for example); with `--cascade=orphan` those stay.
## Errors
[Section titled “Errors”](#errors)
Every error has the same body: a Kubernetes `Status` whose `reason` is a stable error code, plus `fix` (what to do next), `docs` (a page for that code) and `requestId` (quote it to support). The [error reference](/docs/reference/errors/) lists every code.
## Kubernetes tools
[Section titled “Kubernetes tools”](#kubernetes-tools)
The API speaks the Kubernetes wire protocol for these objects, so `kubectl`, `k9s` and client-go work against it: projects are namespaces, `kubectl get jobs.nodus.dev` lists Jobs, and `kubectl explain`, `apply`, `diff`, `wait` and `-w` behave as they do on a cluster. The `nodus` CLI adds what `kubectl` does not have, such as logs, exec and file transfer for Nodus kinds.
# Durable execution
> How an agent run survives restarts, and what it needs from your code to do so.
An agent run is recorded as it goes. Each effect it has is a **step**: creating its sandbox, choosing a model for a prompt, each model call and each command in the sandbox. Nodus stores the result of every step the first time it runs. If Nodus restarts, or a machine fails, the run is picked up again and replays from the top: every step that was recorded returns its stored result, and the run continues with the first step that was not.
## What this guarantees
[Section titled “What this guarantees”](#what-this-guarantees)
* A recorded step never runs again for the same run: no second model call, no second command, no second charge.
* Pure and idempotent steps can retry after an interruption. Nodus model calls reuse their inference key.
* An interrupted sandbox command or model call using your own key can have an uncertain outcome. Its run parks in `Waiting` with reason `NeedsResolution` instead of repeating the effect.
* A run is picked up again by one dispatcher at a time. If two start, one stops without writing.
* A run that reaches a different step than the one recorded at the same position fails with `NonDeterministicReplay`, and its record stays readable.
## What survives a restart
[Section titled “What survives a restart”](#what-survives-a-restart)
The run’s conversation, its answer so far and its status survive, because they come from the recorded steps. The sandbox’s files are not part of the record. If a sandbox is lost, the run creates a new one from the agent’s image and the files are not restored, so keep what must outlive a sandbox in the run’s answer.
## The determinism contract
[Section titled “The determinism contract”](#the-determinism-contract)
Code that runs between steps must give the same result for the same inputs and step results. Clock reads, random numbers and network calls belong inside steps. Agents on Nodus’s Claude follow this for you: the loop that runs the model and its commands is Nodus’s.
## Where the record lives
[Section titled “Where the record lives”](#where-the-record-lives)
Step results are encrypted with your organisation’s key. A run’s recorded payloads expire 30 days after it ends; its status, its cost and its step list stay.
## When an external outcome is uncertain
[Section titled “When an external outcome is uncertain”](#when-an-external-outcome-is-uncertain)
Inspect the run’s steps and the affected files or external system before starting replacement work. A `Started` step means the intent was saved but its outcome was not recorded; it does not prove the command failed or succeeded. Sending another message does not resume this run. Cancel it with `nodus cancel agentrun/` to release its sandbox. The sandbox can continue billing while the run waits. A configured deadline still ends the run.
The managed step-resolution API is not available yet. Do not treat a new run as a safe retry until you have checked the earlier effect.
# Gang networking
> How the members of a distributed Job reach each other across machines and providers, how Nodus checks the path before training starts, and what each path can carry.
Beta
Distributed Jobs (`Job.spec.distributed`) are in beta behind per-org access. Gangs have 2 to 8 members, and `Relayed` gangs have 2 until multi-member Relayed runs are qualified. Ask for access from the console.
A distributed Job runs as a **gang**: one member per machine, every member started together, sometimes on machines from different providers. Before your command starts, Nodus joins the members into a private network that belongs to that gang alone. Each rank gets the addresses of the others in its environment, so `torchrun`, Ray and plain `torch.distributed` work without any networking code of yours.
You choose two things in `spec.distributed`:
* `network`: how close the members must be. `Colocated` keeps them in one provider region, `Regional` allows any provider inside one region class, and `Global` allows anywhere.
* `transport`: which paths you accept. `Direct` (the default) accepts only `Private` and `Direct` paths. `Auto` also accepts `Relayed` paths, which lets members on machines without a kernel network device join.
Nodus picks the path. Every member of a gang uses the same path class, so every rank sees the same addressing and the same NCCL settings.
## Path classes
[Section titled “Path classes”](#path-classes)
| Path | When Nodus uses it | How traffic flows | Encryption | Bandwidth |
| --------- | ----------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------- | ------------------------------------------------ | ------------------------------------------------------------------- |
| `Private` | Every member is in one provider region with a private network | The provider’s private network, opened only between the gang’s members | None added (the provider’s network) | Provider native |
| `Direct` | Every member can create a WireGuard device and the probe finds a direct path between every pair | WireGuard between the members, on a network device named `nodus0` | WireGuard | Measured per pair (see [measured throughput](#measured-throughput)) |
| `Relayed` | A member has no network device of its own (container-only offerings) and `transport: Auto` | WireGuard in user space, carried through Nodus relays | WireGuard end to end; relays see only ciphertext | **Low: small models only** |
### Private
[Section titled “Private”](#private)
Members in the same provider region talk over the provider’s private network. Nodus opens the rendezvous and NCCL ports between the gang’s members only and closes them when the gang ends. The advertised addresses are the members’ private IPs.
### Direct
[Section titled “Direct”](#direct)
Each member runs a WireGuard endpoint next to your container, never inside it: your container gets no extra privileges. The members find each other through the Nodus mesh control and connect peer to peer, punching through NAT where needed. The advertised addresses are the members’ mesh addresses (`100.64.0.0/10`), and NCCL and Gloo use the `nodus0` interface.
`Direct` means direct. If the probe finds that one pair can only connect through a relay, a `transport: Direct` gang is placed again without that pair, and a `transport: Auto` gang continues as `Relayed`. If a running pair loses its direct path, its traffic falls back to a relay without interrupting training, and the Job’s `NetworkDegraded` condition becomes true.
### Relayed
[Section titled “Relayed”](#relayed)
Container-only offerings give your container no network device, no `CAP_NET_ADMIN` and no UDP. Their members can still join a gang: each runs WireGuard in user space and sends peer traffic through two relays that Nodus operates. To reach that WireGuard endpoint, your processes start with a small preloaded library, `libnodus-netshim.so`, which redirects only connections to the other members of your gang. Everything else, such as downloads, datasets, object storage and model APIs, goes out directly as usual.
The advertised addresses are the members’ own `eth0` addresses, so `MASTER_ADDR`, the `torchrun` rendezvous and the listeners NCCL opens on ephemeral ports all work unchanged, in both directions. The relays mesh with each other: if one goes down, members move to the other within seconds and the TCP connections inside WireGuard survive.
## What every rank sees
[Section titled “What every rank sees”](#what-every-rank-sees)
The path decides a few variables; the rest of the distributed environment is the same on every path.
| Variable | `Private` | `Direct` | `Relayed` |
| ---------------------------------------------------------- | --------------------- | -------------- | -------------------------------------------------- |
| `NODUS_GANG_TRANSPORT` | `private` | `direct` | `relayed` |
| `MASTER_ADDR`, `PET_RDZV_ENDPOINT`, `NODUS_NODE_IPS` | Private IPs | Mesh addresses | `eth0` addresses |
| `NCCL_SOCKET_IFNAME`, `GLOO_SOCKET_IFNAME` | The private interface | `nodus0` | `eth0` |
| `LD_PRELOAD`, `NODUS_NETSHIM_PEERS`, `NODUS_NETSHIM_SOCKS` | Not set | Not set | The shim first, then your image’s own `LD_PRELOAD` |
Your own `NCCL_*` values, such as `NCCL_DEBUG=INFO`, win, except `NCCL_NET`, `NCCL_SOCKET_IFNAME`, `NCCL_SOCKET_FAMILY` and `NCCL_IB_DISABLE`, which the path fixes. A spec `LD_PRELOAD` is rejected when `transport: Auto`, because Nodus needs to place the shim first on a `Relayed` path. An `LD_PRELOAD` set in the image, such as jemalloc or tcmalloc, is kept after the shim.
## One network per epoch
[Section titled “One network per epoch”](#one-network-per-epoch)
A gang’s network lives exactly as long as one **epoch** of the gang: one attempt at running it with a fixed set of members. Each epoch gets a fresh mesh identity with single-use join keys that expire after 15 minutes, and its members can reach only each other. When Nodus restarts the gang, for example to replace a lost member, it fences the old epoch first: the old network is torn down, and a member left behind on it can no longer reach the new gang, nor can the new gang reach it. Members of other gangs, in your org or anyone else’s, are never reachable.
## The probe
[Section titled “The probe”](#the-probe)
Before your command starts on any member, every member checks its path to every other member. Nodus bills this time as boot time.
1. **Readiness.** Each member opens three fresh TCP connections to every peer, on ephemeral ports, within 60 seconds, through the same path your processes will use.
2. **Measurement.** For each pair, the member records the round-trip time, a 10-second TCP throughput sample and the path class it actually got: `direct`, or `derp:` when the pair goes through a relay.
3. **Shim self-test** (`Relayed` only). The member runs a check program with your image’s own loader to confirm that the shim loads.
The slowest pair is published in `status.gang.probe` as `tcpGbps` and `rttMs`, with its path, and sets the `NetworkQualified` condition. A pair that cannot connect, or a path worse than your `transport` allows, fails the epoch and Nodus places the gang again; that time is not billed to you. The one exception is an image whose programs cannot load the shim: placing the gang elsewhere cannot fix it, so the Job fails at once with `ShimNotLoaded` and the time is billed (see [caveats](#caveats)).
`tcpGbps` is a TCP throughput sample between two members, not NCCL bus bandwidth. Read it as an upper bound on what one connection between that pair can carry.
```console
$ nodus get job llama-ft -o jsonpath='{.status.gang.probe}'
{"tcpGbps":0.41,"rttMs":38.2,"path":"derp:nodus-us","measuredTime":"…"}
```
## Measured throughput
[Section titled “Measured throughput”](#measured-throughput)
Nodus publishes only measured numbers: the probe’s TCP throughput and RTT, and NCCL all-reduce bus bandwidth from a separate benchmark on 2 × 1 H100. Nodus places a gang only on combinations of offerings it has qualified for the path, and a row appears here once its qualification run passes.
| Path | Members | TCP throughput (probe) | RTT | NCCL bus bandwidth |
| --------- | ----------------------------------- | ---------------------- | ----------------- | ------------------ |
| `Private` | One provider region | Not yet published | Not yet published | Not yet published |
| `Direct` | VMs in different providers | Not yet published | Not yet published | Not yet published |
| `Relayed` | A container-only offering with a VM | Not yet published | Not yet published | Not yet published |
Until the benchmark publishes, treat `Relayed` as low bandwidth: the scheduler assumes it is four times slower than `Direct` when it compares placements, and the estimate warns about it.
## Billing
[Section titled “Billing”](#billing)
`Private` and `Direct` traffic costs nothing beyond the members’ own time. `Relayed` traffic crosses Nodus relays and is billed per GiB on the `mesh-relay` line of your usage, measured on the relays as the bytes each member sends. Mesh traffic never counts as container egress. Each `Relayed` gang may send up to 1 Gbit/s through the relays, split evenly across its members; the relays are shared and best effort within that limit.
## Caveats
[Section titled “Caveats”](#caveats)
* **Synchronous training across providers is bound by the WAN.** Expect cross-provider data-parallel training to be limited by network bandwidth and latency, not by the GPUs. `Relayed` suits small models, algorithms that communicate little, reinforcement learning with separate rollouts, and getting N GPUs now wherever they are.
* **`Relayed` needs dynamically linked glibc programs.** The shim is a preloaded library, so it loads only into programs that use the system’s glibc loader. Launchers built on musl (Alpine images) or linked statically never load it and cannot reach their peers. The probe detects this and the Job fails with `ShimNotLoaded`. Fix it by using a glibc-based image (any Debian, Ubuntu or CUDA image works), or by setting `transport: Direct`.
* **`Relayed` is low bandwidth.** All gang traffic crosses the relays, under the per-gang limit above. Nodus makes no throughput claim for `Relayed` until the benchmark is published.
* **Advertised addresses must be distinct and exclusive.** On `Relayed`, each member is reached at its own `eth0` address, so no two members of a gang may share one, and no two running `Relayed` gangs may route the same address. When placement produces such a collision, Nodus places the members again; it shows as `IPCollision` in the `GangReady` condition while it does.
* **Only connections to peers go through the mesh.** On `Relayed`, connections to other members are redirected; UDP and connections to any other address are not. Tools that need UDP between members, or connect to peers through a hostname that does not resolve to their advertised address, do not work on `Relayed`.
* **The probe measures TCP.** A good `tcpGbps` does not guarantee NCCL performance; NCCL’s numbers come only from the benchmark in the table above.
# Lifecycles
> The phases a Job, Pipeline or Sweep moves through, what moves it, and how spec.state and conditions relate.
Jobs, Pipelines and Sweeps share one lifecycle: they run until they finish. Each shows where it is in `status.phase`, why in `status.reason` and `status.message`, and the details in `status.conditions`. You steer it with `spec.state`.
## Phases
[Section titled “Phases”](#phases)
| Phase | Meaning | Billed |
| -------------- | ----------------------------------------------------------------------------------- | ------------------------------- |
| `Queued` | Admitted; waiting for capacity that fits and for a funded hold | No |
| `Provisioning` | Capacity acquired; the image, source, inputs and any saved state are being prepared | Yes |
| `Running` | The command is running | Yes |
| `Recovering` | The capacity was lost; the work is moving to new capacity | For new capacity, once acquired |
| `Suspending` | State is being saved and compute released | Yes |
| `Suspended` | Paused with its state saved; no compute is held | No compute |
| `Cancelling` | Stopping and releasing compute | Until released |
| `Succeeded` | Finished; outputs collected | No |
| `Failed` | Finished without success; `status.reason` says why | No |
| `Cancelled` | Stopped by `spec.state: Cancelled` or a delete | No |
`Succeeded`, `Failed` and `Cancelled` are final: once there, the phase never changes and nothing more is billed.
## Transitions
[Section titled “Transitions”](#transitions)
| From | To | When |
| ------------------------------------------------- | -------------- | ------------------------------------------------------------------------------------ |
| (new) | `Queued` | The object is admitted |
| `Queued` | `Provisioning` | Capacity is placed and acquired |
| `Queued` | `Failed` | Nothing fits within `placement.queueTimeout` (`CapacityUnavailable`) |
| `Provisioning` | `Running` | The command starts |
| `Provisioning` | `Failed` | The container cannot start (`LaunchFailed`, `ImagePullFailed`) |
| `Running` | `Recovering` | The capacity is lost |
| `Recovering` | `Running` | The work restarts on new capacity |
| `Recovering` | `Failed` | Recovery limits are used up (`RecoveryLimitExceeded`, `NoProgress`, `RestoreFailed`) |
| `Running` | `Succeeded` | The work is done and its outputs are collected |
| `Running` | `Failed` | The command failed more than `backoffLimit` allows, or `timeout` elapsed |
| `Queued`, `Provisioning`, `Running`, `Recovering` | `Suspending` | `spec.state: Suspended`, credits or a budget ran out, or `maxCostUSD` was reached |
| `Suspending` | `Suspended` | State is saved and compute released |
| `Suspending` | `Running` | A suspend you asked for could not save state (`SuspendFailed`) |
| `Suspending` | `Failed` | Saving state failed permanently or timed out |
| `Suspended` | `Queued` | Resumed, and funds cover a new hold |
| any unfinished phase | `Cancelling` | `spec.state: Cancelled`, or the object is deleted |
| `Cancelling` | `Cancelled` | Compute is released and the final charge posted |
## spec.state
[Section titled “spec.state”](#specstate)
`spec.state` is what you want; `status.phase` is what is happening. `nodus suspend`, `nodus resume` and `nodus cancel` set it, and so can a manifest:
| `spec.state` | Effect |
| ------------------- | -------------------------------------------- |
| `Running` (default) | Run, or resume from `Suspended` |
| `Suspended` | Save state, release compute and stop billing |
| `Cancelled` | Stop for good |
A suspend caused by money (credits, a budget or `maxCostUSD`) leaves `spec.state` as it is and resumes on its own once the work is funded again or the cap is raised. Time spent suspended does not count against `timeout`.
## Conditions
[Section titled “Conditions”](#conditions)
| Condition | Meaning |
| ------------------ | ------------------------------------------------------------------------------ |
| `Admitted` | The object passed admission |
| `Scheduled` | Capacity is placed; while false, its message says what it is waiting for |
| `Funded` | Credits and budgets cover the work; false while it is stopped for money |
| `Ready` | The command is running |
| `Checkpointed` | The latest state save succeeded; false with `CheckpointFailed` when it did not |
| `Suspended` | The work is suspended; `SuspendFailed` when a suspend could not save state |
| `OutputsCommitted` | Every declared output is collected |
| `SinksLoaded` | Every output sink has loaded into its table |
Wait on a phase or a condition from the CLI:
```console
$ nodus wait job/train --for=jsonpath='{.status.phase}'=Succeeded --timeout 2h
$ nodus wait job/train --for=condition=OutputsCommitted
```
## Pipelines and Sweeps
[Section titled “Pipelines and Sweeps”](#pipelines-and-sweeps)
A Pipeline or a Sweep takes its phase from its child Jobs. It is `Queued` until its first Job exists, `Running` while any runs, `Suspended` when every unfinished Job is suspended, and final once every Job is final: `Succeeded` if all succeeded, otherwise `Failed` with reason `StageFailed` (Pipeline) or `CellsFailed` (Sweep). Its `spec.state` passes to every unfinished child, and its `maxCostUSD` caps the spend of all its children together.
# Container runtime contract
> The directories, environment variables and sockets your code sees inside every Nodus container.
Every Job, Sandbox and Function worker runs your image in an isolated container. The same contract holds on every offering: your code can rely on the paths and variables below wherever Nodus places it.
## Directories
[Section titled “Directories”](#directories)
| Path | Variable | What it is for |
| --------------------------- | ---------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `/nodus/state` | `NODUS_STATE_DIR` | Recovery state. Write your model, optimizer and progress files here; Nodus checkpoints this directory and restores it on the next attempt. |
| `/nodus/inputs/` | `NODUS_INPUT_` | Each declared input, read-only, downloaded before your command starts. The variable names the file (a URL or one object) or the directory (an object prefix, such as an earlier stage’s outputs). An input may set its own `path` instead. |
| `/nodus/outputs` | `NODUS_OUTPUT_DIR` | Results. Every regular file here is uploaded when your command exits with code 0, with its SHA-256, and appears as the `outputs` output. Symbolic links are not collected. |
| `/dev/shm` | | Shared memory for NCCL between GPUs and for data-loader workers: half the container’s memory limit, or half the machine’s memory without one. It counts against the memory limit. |
| `/etc/nodus/hostfile` | | Multi-node Jobs (Beta): one ` slots=` line per node in rank order, for DeepSpeed and MPI tools. |
| `/run/secrets//` | | Secret values as read-only files (mode `0400`) on a memory-backed filesystem, so they never reach a disk. The environment also carries each value, named by its key. |
| `/run/nodus/events.sock` | `NODUS_EVENTS_SOCKET` | Progress, metrics and the checkpoint handshake (below). |
| `/run/nodus/api.sock` | `NODUS_RUNTIME_SOCKET` | The Nodus API and model calls, authenticated as your Job, Sandbox or Function (below). |
`NODUS_CHECKPOINT_DIR` is kept as another name for `NODUS_STATE_DIR` for existing programs.
## Environment variables
[Section titled “Environment variables”](#environment-variables)
| Variable | Set for | Meaning |
| -------------------------------------------------- | -------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `NODUS_ATTEMPT` | Every container | The attempt id. A retried or recovered run gets a new attempt. |
| `NODUS_JOB`, `NODUS_INDEX`, `JOB_COMPLETION_INDEX` | Jobs | The Job and, for indexed Jobs, this index. |
| `NODUS_RESTORED` | Recovered attempts | `1` when `/nodus/state` was restored from a checkpoint before your command started. |
| `NODUS_CURSOR_COMPLETED`, `NODUS_CURSOR_TOTAL` | Restartable Jobs | The progress cursor your program last reported. |
| `NODUS_PARAM_` | Sweep cells | The cell’s parameters. |
| `PET_NPROC_PER_NODE` | Single-node GPU Jobs | The GPU count, which `torchrun` reads as its `--nproc-per-node` default, so `torchrun train.py` starts one process per GPU. Your own value or flag wins. |
Restoring a checkpoint restores files, never process memory: your program starts from the beginning and reads its own progress files from `/nodus/state`. Nodus never adds resume flags to your command.
## Network
[Section titled “Network”](#network)
Jobs, Functions and Workspaces reach the public Internet by default; Sandboxes and Agents have no outbound network unless their spec opens it. Either way a container never reaches private addresses (`10.0.0.0/8`, `172.16.0.0/12`, `192.168.0.0/16`), cloud metadata services, other containers or the machine it runs on, and has no IPv6. `localhost` always works inside the container. `/etc/resolv.conf` points at public resolvers, and `/etc/hosts` resolves `localhost`.
With an allow list (`egress: AllowList`) the container has no direct route out. `HTTPS_PROXY` and `HTTP_PROXY` point at a proxy on the container’s loopback (`http://127.0.0.1:3128`) that reaches only the listed hosts, on ports 80 and 443. A `*.example.com` entry allows every subdomain of `example.com`. Most HTTP clients (`pip`, `npm`, `curl`, `git` over HTTPS, the Python and Node SDKs) use the proxy on their own. A listed name that resolves to a private address is still refused.
CPU Sandboxes run under gVisor, whose `localhost` belongs to the sandbox alone, so there the proxy and the inference proxy listen on the sandbox’s gateway address instead of `127.0.0.1`. Read the address from `HTTPS_PROXY` and `OPENAI_BASE_URL` rather than writing `127.0.0.1` into your code; `NO_PROXY` already covers it.
On a multi-node Job whose nodes share a private network, each rank’s `eth0` lists the node’s private address first, so NCCL, Gloo and `torchrun` advertise an address the other ranks reach. Ranks reach each other only on the rendezvous and NCCL ports, whatever the Job’s outbound setting.
## GPUs
[Section titled “GPUs”](#gpus)
A container on a GPU offering sees exactly the GPUs assigned to it, numbered from 0 as CUDA and PyTorch see them, with the NVIDIA driver libraries mounted read-only. It never sees another container’s GPUs. Use `nvidia-smi` or `torch.cuda.device_count()` to check what you have; do not set `CUDA_VISIBLE_DEVICES` yourself.
## Sidecars and the init command
[Section titled “Sidecars and the init command”](#sidecars-and-the-init-command)
Sidecars start before your command, in the order you list them, in the same container: they share its files, network, environment and logs. Each must answer its readiness probe (an HTTP `GET` of its path, any status below 400, or a TCP connection to its port) within 15 minutes before the next starts, and they stop when your command exits. A sidecar runs from your container’s image.
`initCommand` runs after the sidecars and before your command, for preflight checks such as imports, free disk or the GPU count. Your command starts only if it exits with code 0; any other code fails the attempt with that code, and running past its timeout (30 minutes unless you set one) fails it with code 124. Both run inside your billed time.
## Commands you run in a container
[Section titled “Commands you run in a container”](#commands-you-run-in-a-container)
Commands you start in a running container (`nodus exec`, Processes) run with the same environment, secrets and directories as your main command, plus any variables you add. Their output is streamed to you and, for Processes, kept with your logs with secret values masked. A command started from a terminal session stops when you disconnect; a Process keeps running until it exits, you cancel it, its timeout passes or the container stops.
Port forwarding and preview URLs reach a server listening on the container’s `localhost`, whatever its outbound network setting. In a CPU Sandbox they reach the sandbox’s own address instead, so listen on all interfaces (`0.0.0.0`); the same holds for a sidecar’s readiness port. File operations (`nodus cp`, the console Files tab) read, write and watch paths as your container sees them, including `/nodus/state` and `/nodus/outputs`, with the permissions of your container’s user: files written this way belong to that user and appear only once complete. `/proc`, `/sys` and `/dev` are not available to them.
## The events socket
[Section titled “The events socket”](#the-events-socket)
The socket speaks newline-delimited JSON, one object of at most 16 KiB per line. Send telemetry as objects with a `type` and, to make retries safe, a unique `id`:
```json
{"id": "evt-10", "type": "progress", "completed": 120, "total": 1000}
{"id": "evt-11", "type": "log.metrics", "step": 1200, "loss": 0.41}
```
When a program cannot reach the socket, it can print the same object on standard output after the prefix `nodus.event `. The line also stays in your logs.
A program that sends no `log.metrics` events still gets its training metrics charted: Nodus reads loss, learning rate, epoch, step and `eval_*` values from Hugging Face Trainer and PyTorch Lightning log lines until the program sends a `log.metrics` event of its own. With `recovery.checkpoint.integration: HFTrainer`, the Trainer helper package is mounted at `/.nodus/python` and put first on `PYTHONPATH`, so its `sitecustomize` registers the Nodus callback in images without the Nodus SDK.
### Checkpoint handshake
[Section titled “Checkpoint handshake”](#checkpoint-handshake)
Nodus decides when to checkpoint. To save a consistent state first, send `{"type": "checkpoint.subscribe"}` once. Before each checkpoint Nodus writes a request on the connection:
```json
{"type": "checkpoint.request", "requestId": "ck-1790000000000", "seq": 1790000000000, "urgent": true}
```
Finish writing your files to `/nodus/state`, then answer with `{"type": "checkpoint.ready", "requestId": "ck-1790000000000"}`. `urgent` means the capacity is about to go away: save quickly. A program that never subscribes is checkpointed without being asked, so write your state files atomically (write to a temporary name, then rename).
A checkpoint of an empty state directory never replaces an earlier checkpoint that had files in it.
## The API socket
[Section titled “The API socket”](#the-api-socket)
`/run/nodus/api.sock` is HTTP over a Unix socket. Requests under `/apis/nodus.dev/` reach the Nodus API, and requests under `/v1/` (OpenAI- and Anthropic-compatible routes) reach Nodus inference. Both carry the service account token of the Job, Sandbox or Function the container belongs to: Nodus adds it to each request, so the token is never in your environment or files, and any `Authorization` header you send is replaced.
```sh
curl --unix-socket "$NODUS_RUNTIME_SOCKET" http://nodus/apis/nodus.dev/v1/...
```
With the inference proxy enabled, the same model routes are also served on `http://127.0.0.1:7777` inside the container (on the gateway address in a CPU Sandbox), even with outbound network off, and `OPENAI_BASE_URL`, `ANTHROPIC_BASE_URL`, `OPENAI_API_KEY` and `ANTHROPIC_API_KEY` are set for it. The keys are placeholders, so SDKs and tools that take a base URL work unchanged. That port serves nothing but model calls.
## Signals and exit codes
[Section titled “Signals and exit codes”](#signals-and-exit-codes)
Your command runs under a small init process (`tini`, at `/.nodus/bin`) that forwards signals to your command’s process group and reaps finished child processes, so your command does not need to be written as PID 1.
A stop sends `SIGTERM` to your command, then `SIGKILL` after the stop grace period (30 s unless your spec sets another). Exit code 0 completes the attempt; any other code, or a signal, fails it with your exit code and the last lines of your logs, with secret values masked.
# Pricing
> How Nodus sets your rate, what each second and token costs, and how to see the price before you launch.
Nodus is prepaid: you buy credits, every run shows its rate before it starts, and the rate stays frozen for that run. The numbers on this page come from the same price list the API uses, so the [pricing reference](/docs/reference/pricing/) and your estimates always agree. This page explains how those numbers are set.
## See the price before you launch
[Section titled “See the price before you launch”](#see-the-price-before-you-launch)
`nodus run --dry-run` and the console review screen show the estimate for a run: the hourly rate, the expected cost range, the expected boot and teardown time as their own lines, and the minimum charge. The estimate is valid for up to 31 minutes. Launching with the estimate’s `ETag` holds you to it: if prices change in between, the launch is refused and you review the new estimate.
## Machines Nodus rents for you
[Section titled “Machines Nodus rents for you”](#machines-nodus-rents-for-you)
A Job, a GPU Sandbox, a GPU Workspace or a GPU Function worker runs on a machine Nodus rents for you. For that machine you pay exactly what the provider bills, divided by 0.875, so Nodus keeps 12.5 % of what you pay:
* **From creation to deletion.** The clock runs from the moment the provider starts billing until the machine is confirmed deleted. Your usage is itemized by segment: `Boot` (start-up, image pull, readiness), `Restore` (resuming from a checkpoint), `Running` and `Teardown` (stop to deletion, plus the provider’s billing increment, charged once per machine).
* **Never above list.** Every accelerator and count has a list price. Nodus never places your work on a machine whose rate would exceed it.
* **Frozen for the run.** The rate is fixed when the machine is acquired. A later price change applies only to runs launched after it takes effect.
Spare machines Nodus starts to finish your work sooner, and failures Nodus causes, are never charged to you.
The list price is the ceiling; the “from” rate is the lowest rate a machine is available at right now, refreshed every minute. Your estimate shows the rate for your run.
| Accelerator | ×1 list | ×1 from | ×8 list | ×8 from |
| ------------- | ------- | --------- | ------- | --------- |
| A10 | $0.89 | List only | $7.12 | List only |
| A100 40G | $1.49 | List only | $11.92 | List only |
| A100 40G PCIE | $1.39 | List only | $11.12 | List only |
| A100 80G | $1.99 | List only | $15.92 | List only |
| A100 80G PCIE | $1.89 | List only | $15.12 | List only |
| B200 | $6.49 | List only | $51.92 | List only |
| H100 PCIE | $2.99 | List only | $23.92 | List only |
| H100 SXM | $3.29 | List only | $26.32 | List only |
| H200 | $4.29 | List only | $34.32 | List only |
| L4 | $0.89 | List only | $7.12 | List only |
| L40S | $1.29 | List only | $10.32 | List only |
| RTX 3090 | $0.59 | List only | $4.72 | List only |
| RTX 4090 | $0.69 | List only | $5.52 | List only |
| RTX 6000 ADA | $1.09 | List only | $8.72 | List only |
| RTX A6000 | $0.79 | List only | $6.32 | List only |
USD per hour for the whole machine. Pricebook 2026.10.8, effective 2026-10-01.
## Sandboxes, Functions and builds
[Section titled “Sandboxes, Functions and builds”](#sandboxes-functions-and-builds)
CPU Sandboxes, Functions, agent workers, image builds and CPU Jobs placed on Nodus nodes are billed per second from placement to release, at published rates per vCPU, per GiB of memory and per GiB of disk above 10 GiB per vCPU. The smallest billable shape is 0.25 vCPU with 512 MiB of memory; smaller requests are billed at that shape.
## Models
[Section titled “Models”](#models)
Each model’s price is the model’s cost divided by 0.95. When a model is served from more than one source, it is listed once, at the price of the cheapest source available now, and a request is charged the price it was accepted at even if another source finishes it. A request is priced once: every input, cached, output, audio or speech unit is added up exactly, then rounded up to the next micro-dollar. A model is served only for the operations it has a price for. If Nodus cannot confirm the outcome of a request, you are not charged for it.
Indra (`nodus/indra`) picks one of 10 models for each request. On the Indra plan you pay the chosen model’s rates plus the routing call (0.044211 USD per 1M input tokens), in one price per request. Without the plan, free Indra routes among the three cheapest models at no charge, up to 100 requests and 200,000 tokens per organization per UTC day.
The Indra plan costs 20.00 USD a month and includes 20.00 USD of `nodus/indra` usage each month at the listed prices. The allowance expires at the end of each month, and usage beyond it draws from your credits.
Agents that run on Nodus-managed models are billed from your credits at the model’s cost divided by 0.875, on the agent run. The plan allowance does not apply to them.
## Storage, egress and credits
[Section titled “Storage, egress and credits”](#storage-egress-and-credits)
* **Storage:** 10 GB per org is included; beyond it, retained bytes are billed per GB-month.
* **Egress:** 10 GiB per org per day is included; beyond it, egress is billed per GiB up to your daily egress quota.
* **New organizations** receive 30.00 USD of credit, valid for 30 days.
## Rounding
[Section titled “Rounding”](#rounding)
Money is counted in micro-dollars. A running machine is charged as it goes, rounded down, and settled when it stops, rounded up once, so charging in many small windows costs the same as charging once. Current rates for every line are in the [pricing reference](/docs/reference/pricing/).
# Sandbox isolation
> What keeps a Sandbox apart from the machine, from other Sandboxes and from the network, and what is kept when it stops.
A [Sandbox](/docs/guides/sandboxes/) runs code you did not write, such as the output of a model, so the boundary around it matters more than it does for your own jobs. This page lists the layers of that boundary, in the order a piece of code meets them.
## Where Sandboxes run
[Section titled “Where Sandboxes run”](#where-sandboxes-run)
Sandboxes run on **CPU machines Nodus operates**, not on GPU machines rented for your jobs. Nodus packs machines **per organization**: a machine serves one organization at a time, your Sandboxes share it with your other Sandboxes, agent workers and CPU jobs, and with nobody else’s. An empty machine is destroyed after 10 minutes and never handed to another organization, so one organization’s files and memory never sit on a machine another organization uses next.
## Layers
[Section titled “Layers”](#layers)
### 1. A user-space kernel (gVisor)
[Section titled “1. A user-space kernel (gVisor)”](#1-a-user-space-kernel-gvisor)
Each Sandbox runs under [gVisor](https://gvisor.dev/) (`runsc`, in its `systrap` mode). gVisor answers the Sandbox’s system calls in its own kernel written in a memory-safe language and passes the host kernel only a small, filtered set of calls. A bug in the host kernel’s handling of an unusual system call is reachable from a Sandbox only through that narrow filter, not directly.
### 2. A network namespace and a firewall per Sandbox
[Section titled “2. A network namespace and a firewall per Sandbox”](#2-a-network-namespace-and-a-firewall-per-sandbox)
Every Sandbox gets its own network namespace with a default-deny `nftables` firewall:
* Private ranges (RFC 1918), carrier-grade NAT, link-local addresses and cloud metadata addresses are always blocked, whatever the policy, so a Sandbox cannot reach the machine, its neighbours or the cloud’s metadata service.
* Policy `Deny` (the default) has no route out at all.
* Policy `Open` translates traffic to public addresses only.
Nothing is allowed in from the network. Bytes in and out are counted, and a blocked connection raises an `EgressDenied` Event on the Sandbox.
### 3. Resource limits
[Section titled “3. Resource limits”](#3-resource-limits)
Each Sandbox runs in its own cgroup (v2) that limits CPU and memory to what you asked for and the number of processes to 256, so a fork bomb or a memory leak stops at the Sandbox’s own limits instead of slowing its neighbours. The root filesystem is backed by a per-Sandbox file under a disk quota (`resources.disk`), not by a directory shared with other Sandboxes.
### 4. An unprivileged process
[Section titled “4. An unprivileged process”](#4-an-unprivileged-process)
Commands run as a non-root user with no Linux capabilities and `no_new_privs` set, so the process cannot gain privileges through a set-uid program. The user comes from the image. The Nodus helper that serves commands and files runs outside the user’s process tree and is read-only to it.
## Continuity: what survives a stop
[Section titled “Continuity: what survives a stop”](#continuity-what-survives-a-stop)
`spec.continuity.mode` decides what a Sandbox keeps when it stops or when its machine is lost:
| Mode | A stop keeps | If the machine is lost |
| ----------------------- | ---------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------ |
| `Snapshotted` (default) | The filesystem changes the Sandbox made, `/workspace` first among them | The Sandbox restarts on another machine from its last snapshot; work since that snapshot is gone |
| `Ephemeral` | Nothing: the next start is empty | The Sandbox moves to `Failed` with the reason `NodeLost` |
Processes never survive a stop: a start runs a fresh container from the image with the saved files restored, so start long-lived servers again from your code. Keep your work in `/workspace`, the directory the Sandbox starts in. `Snapshotted` is the right choice for an agent that works in a repository; `Ephemeral` suits a fresh sandbox per task. Snapshots are taken at every stop and about every 10 minutes while the Sandbox runs.
## Where gVisor is not available
[Section titled “Where gVisor is not available”](#where-gvisor-is-not-available)
Production Nodus CPU nodes ship with `runsc`, and Sandboxes there run under gVisor as described above. On a node without `runsc` (the local development provider on a laptop, or a Linux host where gVisor is not installed), `nodusd` runs the Sandbox under `runc` instead, with the same network namespace, firewall, cgroup limits and unprivileged process. Layers 2 to 4 are identical; layer 1 is not: a `runc` Sandbox shares the host’s Linux kernel, so it is a boundary for development and trusted code, not for hostile code.
Caution
Do not run untrusted code on a development node that lacks `runsc`.
## What a Sandbox does not protect against
[Section titled “What a Sandbox does not protect against”](#what-a-sandbox-does-not-protect-against)
* **Secrets you put in it.** Anything in a Sandbox’s environment or files is readable by the code running there. With `egress.policy: Open`, that code can send it anywhere on the internet. Keep `Deny` unless the work needs the network, and give a Sandbox only the credentials its task needs.
* **Resource use up to your limits.** A Sandbox can use all the CPU and memory it asked for, and is billed for them while it runs. Set `maxCostUSD` to cap the spend of code you do not control.
# Placement and scheduling profiles
> How Nodus chooses where a run executes, what the estimate includes, and how profiles, deadlines and budgets change the choice.
By default Nodus places each run on the **cheapest offering that can start it now**: the lowest hourly rate among the offerings that fit your request, have a machine free and stay within your limits. If that offering cannot be used, the next cheapest takes the run, and so on. You pay the rate of the offering the run starts on, shown in the estimate and frozen for the run. This page explains what the estimate includes and the settings that change the choice.
## Cost to completion
[Section titled “Cost to completion”](#cost-to-completion)
For every offering that fits your request, Nodus estimates the whole bill of the run. The estimate shows it for the chosen offering, and the `Cost` profile ranks offerings on it:
* **Startup**: the machine’s boot, the image pull and any restore from a checkpoint. You pay for these because the capacity bills from the moment it is created.
* **Running time**: from your declared `expectedDuration`, a training runtime’s measured estimate, or the history of earlier runs with the same labels on the same accelerator family.
* **Teardown**: the time until the machine is confirmed deleted.
* **Rounding**: each machine bills in increments, and the estimate rounds up the same way.
* **Interruptions**: for interruptible capacity, the expected number of reclaims times the work redone after each one. Work that checkpoints loses only the stretch since its last save; work that saves nothing loses half its run on average.
Under `Cost`, a cheap interruptible offering therefore wins only when its expected cost, lost work included, still beats on-demand capacity. A run whose checkpoints are all empty is costed as if it saves nothing.
When the running time is unknown, every profile ranks offerings by hourly rate and the estimate shows the rate and the startup and teardown charge but no total.
## Profiles
[Section titled “Profiles”](#profiles)
`placement.profile` picks how the scheduler trades cost against time:
| Profile | Chooses |
| -------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `Balanced` (default) | The lowest hourly rate that meets your deadline and budget. Among offerings within 2 % of that rate, the healthier one, or the one that starts sooner, fits better or is already warm |
| `Cost` | The lowest expected cost to completion, startup, teardown and lost work included, even if it is slower to start. Among offerings within 2 % of that cost, warm capacity and a better fit win |
| `Speed` | The fastest expected finish among offerings within 1.5 × the cheapest expected cost |
`Balanced` never chooses an offering whose rate is more than 2 % above the cheapest that meets your limits, and `Cost` never chooses one whose expected cost is more than 2 % above the cheapest at equal health; `interruptible: Prefer` gives interruptible capacity a 10 % allowance. Under `Cost` and `Speed`, an offering that often fails to start counts as dearer by that risk; under every profile, one that recently failed to start ranks lower for a few minutes within the 2 % band. When two offerings are within 2 % of each other, identical requests are spread across both instead of all taking the same one.
## Deadlines and budgets
[Section titled “Deadlines and budgets”](#deadlines-and-budgets)
* `placement.completeByTime`: offerings whose p90 finish is later are not used.
* `maxCostUSD`, Budgets and your balance: offerings whose expected cost exceeds the money left are not used.
* `placement.maxRateUSDPerHour`: offerings above this rate are not used.
If nothing remains, the run waits in `Queued` and the estimate says why, for example `MissesDeadline` or `ExceedsRemainingBudget`.
## Estimates and the If-Match ceiling
[Section titled “Estimates and the If-Match ceiling”](#estimates-and-the-if-match-ceiling)
`--dry-run=server -o estimate` returns the expected cost p50 and p90, the startup time (cold, and warm when idle capacity of yours fits), the first hold, the minimum charge and `validUntil`, which is at most 31 minutes away. Creating with the estimate’s `If-Match` binds the launch to it: Nodus then uses no offering above the estimated rate plus 10 %, unless you set `placement.maxRateUSDPerHour` yourself. If prices moved beyond that, the run waits with `PriceAboveEstimate` instead of costing more than you saw.
## Reuse before new capacity
[Section titled “Reuse before new capacity”](#reuse-before-new-capacity)
Before buying new capacity, the scheduler considers capacity you already pay for: an idle machine of yours that fits reuses the increment already paid, and CPU work packs onto your Nodus nodes. A [BYOC pool](/docs/guides/pools/) named in `placement.pool` is used first at no hourly charge.
## Why a placement was made
[Section titled “Why a placement was made”](#why-a-placement-was-made)
`nodus describe` shows each attempt’s placement: the profile, the scores of the chosen offering (fit, time to result, cost to complete, recovery value, health), the fallbacks in order and every rejected offering with its reason. Offerings are shown by name, such as `h100-sxm-80g-x8-us`.
## Multi-node runs (Beta)
[Section titled “Multi-node runs (Beta)”](#multi-node-runs-beta)
For `spec.distributed`, Nodus first resolves the topology (nodes × GPUs per node) and then places every node together:
* `network: Colocated` (the default) keeps every node in one location on one network; `Regional` keeps them in one region class; `Global` allows anywhere.
* `transport: Direct` uses only private or direct paths between nodes; `Auto` also allows the relayed mesh, with lower bandwidth.
* The whole gang is priced together, from the slowest node’s start. The estimate also shows the assembly bound: the most a failed assembly can cost.
If no set of offerings satisfies these rules, the run waits with `GangInfeasible`. See [multi-node training](/docs/guides/multi-node/).
# Install the CLI
> Install the nodus CLI on macOS, Linux or Windows, verify the download, turn on shell completion and sign in.
The `nodus` CLI is one self-contained binary for macOS, Linux and Windows on x86-64 and ARM64. Pick one way to install it.
* pip
```sh
pip install nodus-compute
```
The Python package (Python 3.10 or newer) includes the CLI, so `nodus` is on your `PATH` wherever the package is installed. Use this if you also want the Python SDK.
* Homebrew
```sh
brew install --cask nodus-compute/tap/nodus
```
* macOS and Linux
```sh
curl -fsSL https://nodus-compute.ai/install | sh
```
The script installs to `~/.local/bin` and checks the archive against the release’s `checksums.txt` first. `NODUS_VERSION=1.2.3` pins a release and `NODUS_INSTALL_DIR` picks another directory.
* Windows
```powershell
irm https://nodus-compute.ai/install.ps1 | iex
```
Check that it works:
```console
$ nodus version
nodus v1.0.0
```
## Sign in
[Section titled “Sign in”](#sign-in)
```sh
nodus login
```
Your browser opens to confirm the sign-in. If you belong to several orgs, pick the ones this machine should use: the CLI stores one API key per org in the OS keychain and creates one **context** per org. Switch between them with `nodus config use-context `, or pass `--org ` to a single command.
| Where you are | Command |
| --------------------------------------------- | -------------------------------------------------------------------------------- |
| A laptop with a browser | `nodus login` |
| An SSH session or a machine without a browser | `nodus login --device`, then approve the code from any device |
| CI, with a key in a secret | `echo "$NODUS_API_KEY" \| nodus login --with-token`, or just set `NODUS_API_KEY` |
`nodus whoami` shows who you are signed in as, your role, the current project and your available credit. `nodus logout` revokes this machine’s key and removes the context.
Environment variables
The CLI reads six variables, all optional: `NODUS_API_KEY`, `NODUS_API_URL`, `NODUS_ORG`, `NODUS_PROJECT`, `NODUS_CONTEXT` and `NODUS_CONFIG`. A variable wins over the config file for that invocation.
## Shell completion
[Section titled “Shell completion”](#shell-completion)
```sh
nodus completion zsh > "${fpath[1]}/_nodus" # zsh
nodus completion bash > /etc/bash_completion.d/nodus # bash (or ~/.local/share/bash-completion/completions/nodus)
nodus completion fish > ~/.config/fish/completions/nodus.fish
```
Completion covers every command and flag.
## Verify a download
[Section titled “Verify a download”](#verify-a-download)
Every release publishes `checksums.txt`, a keyless cosign signature over it, and an SBOM per archive. To verify an archive you downloaded yourself:
```sh
cosign verify-blob checksums.txt \
--signature checksums.txt.sig --certificate checksums.txt.pem \
--certificate-identity-regexp '^https://github.com/nodus-compute/nodus-platform/' \
--certificate-oidc-issuer https://token.actions.githubusercontent.com
shasum -a 256 --ignore-missing -c checksums.txt
```
## Where the CLI keeps things
[Section titled “Where the CLI keeps things”](#where-the-cli-keeps-things)
| Path | What it holds |
| -------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------- |
| `~/.nodus/config` | Contexts: API server, org and project. It is a kubeconfig, so `KUBECONFIG=~/.nodus/config kubectl get jobs.nodus.dev` works too |
| OS keychain (`~/.nodus/credentials`, mode 0600, where there is none) | One API key per context |
| `~/.nodus/cache/` | The cached list of resource kinds; `nodus api-resources` refreshes it |
If you used Nodus before 1.0, its `~/.nodus/config.toml` is renamed to `config.0x.bak` the first time the new CLI runs; sign in again with `nodus login`. A 0.x `nodus.toml` converts into a manifest you can review and apply:
```sh
nodus convert nodus.toml > job.yaml # fields that do not carry over are listed on stderr
nodus apply -f job.yaml --dry-run=server -o estimate
```
## Upgrade and uninstall
[Section titled “Upgrade and uninstall”](#upgrade-and-uninstall)
Upgrade the same way you installed (`pip install -U nodus-compute`, `brew upgrade --cask nodus`, or run the install script again). To uninstall, run `nodus logout`, remove the binary, and delete `~/.nodus`.
## Next steps
[Section titled “Next steps”](#next-steps)
* [Quickstart](/docs/getting-started/quickstart/): run your own code on a GPU and download its results.
* [Nodus for kubectl users](/docs/getting-started/kubectl-users/): the verbs you already know.
* [CLI reference](/docs/reference/cli/): every command and flag.
# Nodus for kubectl users
> How Nodus maps onto the Kubernetes API model, which kubectl verbs nodus supports, and how to point kubectl itself at Nodus.
If you know kubectl, you already know most of `nodus`. Every Nodus resource (Jobs, Sandboxes, Volumes, Secrets, InferenceEndpoints, Budgets and the rest) is a Kubernetes-style object in the `nodus.dev` API group, with `metadata`, `spec` and `status`, and the CLI speaks the same verbs over all of them.
## The mapping
[Section titled “The mapping”](#the-mapping)
| Kubernetes | Nodus |
| -------------------- | ------------------------------------------------------------------------------------ |
| Cluster | The Nodus API (`https://api.nodus-compute.ai`) |
| Namespace (`-n`) | Project (`-p`; `-n` is accepted) |
| User and credentials | An API key per org, kept in the OS keychain |
| kubeconfig context | One context per org, in `~/.nodus/config` |
| `kubectl get pods` | `nodus get jobs`, `nodus get sandboxes`, `nodus get all` |
| `kubectl exec` | `nodus exec` into a Job, Sandbox or Workspace; each command is recorded as a Process |
Your org comes from the API key, so there is no org in a manifest. Projects are created in the console or with `nodus create -f`, and `default` always exists.
## Verbs you already know
[Section titled “Verbs you already know”](#verbs-you-already-know)
```sh
nodus get jobs -o wide # tables are rendered by the server
nodus get jobs -l team=nlp --field-selector status.phase=Running
nodus get job/train -o jsonpath='{.status.phase}'
nodus get jobs -o custom-columns=NAME:.metadata.name,GPU:.spec.resources.gpu
nodus get jobs -w # watch
nodus describe job/train # includes a Placement section and events
nodus apply -f job.yaml # client-side three-way merge
nodus diff -f job.yaml # what apply would change (exit 1 on differences)
nodus apply -f jobs/ --prune -l app=nightly # delete what was applied before and is gone now
nodus edit job/train # $EDITOR, saved as a merge patch
nodus patch job/train --type merge --patch '{"spec":{"maxCostUSD":"50"}}'
nodus label job/train tier=gold
nodus wait job/train --for=jsonpath='{.status.phase}'=Succeeded --timeout 30m
nodus delete job/train --wait
nodus logs -f job/train
nodus exec -it sb/dev -- bash
nodus shell sandbox/dev # create the Sandbox if needed, then exec -it bash
nodus events -w --for job/train # events as they are recorded
nodus port-forward job/train 6006
nodus explain job.spec.resources
nodus api-resources
nodus auth can-i create jobs
```
Short names work as in kubectl: `sb` for sandboxes, `vol` for volumes, `sec` for secrets, `sa` for service accounts. `nodus api-resources` lists them all.
## What is different
[Section titled “What is different”](#what-is-different)
* **Server dry-run returns a cost estimate.** `nodus apply -f job.yaml --dry-run=server -o estimate` prints the estimated cost, start time and the hold that would be placed, without creating anything.
* **State verbs instead of scaling.** `suspend`, `resume` and `cancel` apply to Jobs, Pipelines and Sweeps; `start` and `stop` to Sandboxes, Workspaces, Functions, Agents and InferenceEndpoints. They set `spec.state`.
* **Requests are annotations.** `nodus request restart sb/dev` (or `nodus rollout restart sb/dev`) asks a controller to act once, recorded on the object.
* **Typed generators.** `nodus create sandbox dev --cpu 2`, `nodus create secret hf --from-literal HF_TOKEN=…`, `nodus create volume weights --size 200Gi`, `nodus create budget research --limit 2000`. Add `--dry-run=client -o yaml` to print the manifest instead of creating it. Kinds that show a secret once print it alone on standard output: `nodus create apikey ci --scopes jobs:write`, `nodus create enrollmenttoken --pool lab` (the host installer) and `nodus create token sa/ci`. `nodus create sshkey --from-file ~/.ssh/id_ed25519.pub` adds a public key, or `--generate` makes the pair.
* **Outputs are files you download.** `nodus cp job/train:outputs/model ./model` copies a declared output and checks its SHA-256 digest.
* **Merge patches only.** Strategic merge patch and server-side apply are not supported; `apply` does the three-way merge in the client and records the applied manifest in a last-applied annotation on the object.
* **Exit codes.** Every command exits `0`, `1` on an error or `2` on a usage error, like kubectl. `nodus run` passes your command’s exit code through instead.
## Use kubectl itself
[Section titled “Use kubectl itself”](#use-kubectl-itself)
`~/.nodus/config` is a real kubeconfig whose user entry runs `nodus auth token` as an exec credential plugin, so kubectl and any client-go tool (k9s included) can read and write Nodus resources:
```sh
export KUBECONFIG=~/.nodus/config
kubectl get jobs.nodus.dev
kubectl apply -f job.yaml
kubectl get jobs.nodus.dev -w
kubectl wait jobs.nodus.dev/train --for=jsonpath='{.status.phase}'=Succeeded
```
Use the full `jobs.nodus.dev` resource name with kubectl, because it also knows the built-in `batch/v1` Jobs.
What kubectl cannot do against Nodus
kubectl’s `logs`, `exec` and `port-forward` only work on Pods, so use `nodus` for those. There are no Pods, Namespaces or other core Kubernetes resources, and `kubectl auth can-i` is not served: use `nodus auth can-i` instead.
## Plugins
[Section titled “Plugins”](#plugins)
Like kubectl, `nodus` runs any executable called `nodus-` on your `PATH` as `nodus `, so you can add your own commands.
# Nodus for Modal users
> Move a Modal app to Nodus. The Python SDK keeps Modal's names wherever the concept matches, so most code changes by one import.
The Nodus Python SDK keeps Modal’s names wherever the concept is the same: `App`, `@app.function`, `.remote()`, `.map()`, `.spawn()`, `@app.cls` with `enter` and `exit` hooks, `Image`, `Volume`, `Secret` and `Sandbox`. Most apps move by changing `import modal` to `import nodus`. This page lists what is the same, what is spelled differently and what Nodus adds.
## Sign in and run
[Section titled “Sign in and run”](#sign-in-and-run)
```console
pip install nodus-compute
nodus login # opens the console, stores a key for each org you pick
nodus run app.py # like `modal run`: an ephemeral App, deleted when the entrypoint returns
nodus deploy app.py # like `modal deploy`: a persistent App
nodus serve app.py # like `modal serve`: redeploys when a file changes
```
In CI, set `NODUS_API_KEY` instead of running `nodus login`.
examples/python/quickstart/app.py
```python
"""Quickstart: one Function called three ways.
Run it with `nodus run examples/python/quickstart/app.py --n 10`.
"""
import nodus
app = nodus.App("quickstart")
@app.function(cpu=1, memory="1Gi", max_cost=1)
def square(x: int) -> int:
return x * x
@app.local_entrypoint()
def main(n: int = 10) -> None:
print("remote:", square.remote(7)) # one call; blocks for the result
call = square.spawn(8) # start without waiting
print("spawned:", call.get(timeout=600))
print("map:", list(square.map(range(n)))) # one call per input, results in input order
```
## Side by side
[Section titled “Side by side”](#side-by-side)
| Modal | Nodus |
| ------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------- |
| `import modal` | `import nodus` |
| `app = modal.App("x")` | `app = nodus.App("x")` |
| `@app.function(gpu="H100", timeout=3600)` | `@app.function(gpu="H100", timeout="1h")` (numbers are also seconds) |
| `f.remote(x)`, `f.map(xs)`, `f.spawn(x)`, `call.get()` | The same |
| `f.local(x)` | The same |
| `@app.local_entrypoint()` | The same |
| `modal.Function.from_name("app", "f")` | `nodus.Function.from_name("app", "f")` (`Function.lookup` also works) |
| `@app.cls()` with `@modal.enter()`, `@modal.method()`, `@modal.exit()` | `@app.cls()` with `@nodus.enter()`, `@nodus.method()`, `@nodus.exit()` |
| `modal.Image.debian_slim().pip_install("torch")` | `nodus.Image.debian_slim().pip_install("torch")` |
| `modal.Image.from_registry(...)`, `.apt_install`, `.run_commands`, `.env`, `.add_local_dir` | The same |
| `modal.Volume.from_name("v", create_if_missing=True)` | The same; `vol.commit()` and `vol.reload()` behave as in Modal |
| `modal.Secret.from_name("hf")`, `Secret.from_dict({...})` | The same, plus `Secret.from_dotenv(".env")` |
| `modal.Sandbox.create(app=app, image=...)` | `nodus.Sandbox.create(image=...)` (no App needed) |
| `sb.exec("python", "-c", "...")`, `p.stdout.read()`, `p.wait()` | The same |
| `sb.open(path, "w")` | The same |
| `sb.tunnels()` | Returns a list of `Tunnel(port, url, public)`; `sb.tunnels.open(8080)` is not available yet |
| `sb.snapshot_filesystem()` | The same (Beta) |
| `@modal.experimental.clustered(size=2)` | `@nodus.clustered(size=2)` (Beta) |
| `min_containers`, `max_containers`, `scaledown_window` | The same, or `min_workers` and `max_workers` |
| `await f.remote.aio(x)` | The same: every blocking call has an `.aio` form |
GPU strings use Modal’s spellings: `"H100"`, `"H100:2"`, `"A100-80GB"`, `"A10G"`, `"L40S"`. A family such as `H100` matches any of its variants; `"H100!"` pins the exact variant. A list such as `["H100", "H200"]` accepts either.
## What Nodus adds
[Section titled “What Nodus adds”](#what-nodus-adds)
Nodus places every call on the cheapest capacity that finishes it on time, and it stops work before money runs out. These arguments have no Modal equivalent:
```python
@app.function(
gpu="H100",
max_cost=40, # a hard cap in USD across this Function's workers
checkpoint="/nodus/state", # files here are saved and restored if capacity is reclaimed
interruptible=True, # allow cheaper interruptible capacity; progress is kept through the checkpoint
region=["us", "eu"], # region classes, not provider regions
)
def train(lr: float) -> dict: ...
print(train.estimate(3e-4)) # dry-run: expected cost, cold and warm start, the hold it needs
```
* `f.estimate(...)` returns the expected cost and start time of one call before you run it.
* A cold start on a GPU with no warm worker prints its expected wait, so a long first call is not a surprise.
* Errors are typed: `nodus.errors.InsufficientCredits` states the amount needed and how to add credit.
* `nodus.Job`, `nodus.Workspace` and `nodus.llm` cover batch jobs, development machines and inference with the same credentials.
## Differences to know
[Section titled “Differences to know”](#differences-to-know)
* **Parametrized classes** (`modal.parameter()`) are not supported. Configure the class in its `@nodus.enter()` hook instead. Calling `MyClass(arg=...)` raises `nodus.errors.Unsupported`.
* **The Python minor version** of the image must equal yours, as in Modal. `Image.debian_slim()` defaults to your version; a mismatch is refused before anything runs.
* **Web endpoints** (`@modal.web_endpoint`, `@modal.asgi_app`) are not available. Use `nodus.InferenceEndpoint` for model serving.
* **`modal.Dict` and `modal.Queue`** have no equivalent. Pass data through return values, a Volume or your own database.
* **Clustered Functions** (Beta) run each `.remote()` or `.spawn()` as one gang; `.map()` over a clustered Function raises `nodus.errors.Unsupported`.
* **Timeouts and durations** accept Go-style strings (`"90s"`, `"6h"`) as well as seconds. Days are not a unit.
* **Money** is always a decimal amount in USD (`max_cost=40` or `"40.00"`).
## Next steps
[Section titled “Next steps”](#next-steps)
* [Python SDK guide](/docs/guides/python/)
* [Functions and classes](/docs/guides/python/functions/)
* [Sandboxes](/docs/guides/python/sandboxes/)
# Run your own code
> Run a training script from your own directory on a GPU with nodus run, follow it, download its output and read what it cost.
This page runs a script from your own directory on a GPU, then downloads the file it wrote. It assumes you have [installed the CLI and signed in](/docs/getting-started/install/). Every command here runs in CI as the `cli/quickstart` example.
1. **Put your code in a directory.** Any directory works; this one holds a small PyTorch training loop that saves its final metrics to `/nodus/outputs/metrics.json`.
train.py
```python
"""Fit a small linear model, print progress and save metrics.json as the declared output "metrics"."""
import json
import pathlib
import torch
device = "cuda" if torch.cuda.is_available() else "cpu"
print("device:", torch.cuda.get_device_name() if device == "cuda" else "cpu")
torch.manual_seed(0)
x = torch.randn(1024, 16, device=device)
y = x @ torch.randn(16, 1, device=device)
w = torch.zeros(16, 1, device=device, requires_grad=True)
opt = torch.optim.SGD([w], lr=0.1)
for step in range(200):
loss = ((x @ w - y) ** 2).mean()
opt.zero_grad()
loss.backward()
opt.step()
if step % 50 == 0:
print(f"step {step} loss {loss.item():.4f}")
print(f"final loss {loss.item():.6f}")
out = pathlib.Path("/nodus/outputs")
out.mkdir(parents=True, exist_ok=True)
(out / "metrics.json").write_text(json.dumps({"final_loss": loss.item(), "steps": 200}))
```
2. **Run it.** From that directory:
```sh
nodus run --name cli-quickstart --gpu L4 --image nodus/pytorch \
--output metrics=/nodus/outputs/metrics.json -- python train.py
```
```console
job/cli-quickstart created · est. $0.01–0.03 · starts in ~2–4 min (cold) · hold $0.50 · Ctrl+C to cancel, -d to detach
✓ Scheduled l4-24g-x1-us · $0.52/h (rate frozen)
✓ Provisioning 1m48s
▶ Running
device: NVIDIA L4
step 0 loss 15.8732
...
final loss 0.000001
✓ Succeeded in 2m31s · $0.02 (boot $0.01 · running $0.01) · kept warm 60 s · nodus describe job/cli-quickstart
```
Before anything is charged you see the estimate and a hold: the most this run can cost before you are asked. `nodus run` then uploads the directory (files listed in `.gitignore` or `.nodusignore` stay local), prints each phase, streams your program’s output on stdout and ends with the cost line. The CLI exits with your command’s exit code, so it works in scripts and CI.
3. **Get the output.** The `--output` flag declared `metrics`; copy it back:
```sh
nodus cp job/cli-quickstart:outputs/metrics metrics.json
```
The file is checked against its SHA-256 digest and only written when it matches.
## Leave it running
[Section titled “Leave it running”](#leave-it-running)
Press **Ctrl+C** during a run and the CLI asks whether to cancel the Job; answer `n` to detach and leave it running. Or start with `-d` to return as soon as the Job exists. A detached Job keeps running:
```sh
nodus get jobs # every Job in the project, with phase and cost so far
nodus get jobs --mine -w # only yours, updating live
nodus logs -f job/cli-quickstart # stream the output again
nodus describe job/cli-quickstart # attempts, placement, cost, conditions and events
nodus cancel job/cli-quickstart # stop it; you pay only for what ran
```
## Control the cost
[Section titled “Control the cost”](#control-the-cost)
| Flag | Effect |
| -------------- | ------------------------------------------------------ |
| `--max-cost 5` | The Job stops gracefully before it spends more than $5 |
| `--timeout 2h` | The Job is stopped after two hours of wall-clock time |
| `--dry-run` | Print the estimate and exit without creating anything |
`nodus billing` shows your balance and this month’s spend; `nodus billing usage --group-by day` breaks it down.
Exit codes
`nodus run` exits with your command’s own exit code. It exits `125` when Nodus could not run the command (an API error or an invalid request), `124` on `--timeout` and `130` when you interrupt it.
## Clean up
[Section titled “Clean up”](#clean-up)
Finished Jobs are kept for 30 days so you can read their logs and outputs, then deleted. To delete one now:
```sh
nodus delete job/cli-quickstart
```
## Next steps
[Section titled “Next steps”](#next-steps)
* [Nodus for kubectl users](/docs/getting-started/kubectl-users/): manifests, `apply`, `get -o`, `wait` and `diff`.
* [CLI reference](/docs/reference/cli/): every command and flag, with tested examples.
# API keys and scopes
> Create API keys for scripts and CI, limit what they can do with scopes and projects, and revoke them.
An API key lets a script, a notebook or CI call Nodus as you. Keys look like `nodus_sk_live_…` and are sent as `Authorization: Bearer `, or through `NODUS_API_KEY`.
```bash
nodus create apikey notebook
```
The key is printed **once**, alone on standard output, so `KEY=$(nodus create apikey notebook)` captures it. Nodus stores only a keyed hash of it, so a lost key cannot be shown again: delete it and create a new one. `-o json` or `-o yaml` prints the whole object, key included.
## Scopes
[Section titled “Scopes”](#scopes)
A key’s scopes limit what it can do. The default, `*`, means “whatever my role allows”, and it follows your role: if your role changes, the key changes with it.
```bash
nodus create apikey ci --scopes jobs:write,volumes:read --projects research --expires 720h
nodus auth can-i create sandboxes # run with the key to check it
```
Scopes are `:read` or `:write` (write includes read), plus `*:read` for read-only access to everything. A key can never hold more than its creator’s role: asking for more is refused with `403` naming the scopes you lack. A key created with another key (or by a connected agent) can hold only what that credential holds, so it cannot ask for `*` or `*:read` unless the creating credential has them itself.
`--scopes` and `--projects` take comma-separated lists, and `--description` says what the key is for. To mint a key for CI that survives its creator, bind it to a service account of the project you work in with `--service-account NAME`.
A key restricted with `--projects` reaches only those projects, and it cannot change org-level settings. If every project it was restricted to is deleted, it reaches no project at all. A key bound to a service account can only be restricted to that service account’s project.
## List and revoke
[Section titled “List and revoke”](#list-and-revoke)
```bash
nodus get apikeys
nodus delete apikey ci
```
A deleted key stops working within 30 seconds everywhere. Listing shows the key’s prefix (`nodus_sk_live_01j9…`), scopes, owner, last use and expiry, never the key itself. Members can delete their own keys. Deleting another member’s key takes an Admin or Owner, and deleting a service account’s key takes permission to manage service accounts in its project.
## Keys the CLI creates
[Section titled “Keys the CLI creates”](#keys-the-cli-creates)
`nodus login` creates one key per org, named `cli--`, with every scope your role grants, valid for 90 days and labelled as launched by the CLI. `nodus logout` revokes it.
## If a key leaks
[Section titled “If a key leaks”](#if-a-key-leaks)
Delete it. Keys are registered with GitHub secret scanning, so a key pushed to a public repository is reported to us and revoked.
# Connected agents
> See and disconnect the MCP clients you let act on Nodus for you.
When an MCP client such as Claude Code, Cursor or Codex connects to Nodus, the consent page asks you to pick an org and the scopes the client may use. Nodus records that choice as an **OAuth grant**. The client can do only what the grant allows, and never more than your own role in that org. It cannot create API keys or other grants.
A client is connected to one org at a time. Approving it again, for any org, replaces the earlier grant.
## Seeing what is connected
[Section titled “Seeing what is connected”](#seeing-what-is-connected)
Console › Settings › Connected agents lists each client with its scopes and when it was last used. Or:
```bash
nodus get oauthgrants
```
You see your own grants. Admins and Owners see every member’s grants in the org.
## Disconnecting
[Section titled “Disconnecting”](#disconnecting)
```bash
nodus delete oauthgrant claude-code-3fa2c1
```
The client stops working within 30 seconds, even with an access token it already holds. A grant also lapses 30 days after its last use, and removing a member from the org disconnects all of their clients there.
# Device login
> Sign the CLI in on a machine without a browser, or from CI with an existing key.
On a remote server or container with no browser:
```bash
nodus login --device
```
The CLI prints a code like `BDWP-HMTR` and a link. Open the link on any device where you are signed in, check that the code matches, pick the orgs and approve. The CLI finishes on its own within a few seconds. The code works for 10 minutes; denying it stops the login.
The CLI polls every 5 seconds. Every key it receives is the same kind `nodus login` creates: one per chosen org, 90 days, revocable with `nodus logout` or `nodus delete apikey`.
## In CI
[Section titled “In CI”](#in-ci)
Pass an existing key instead of signing in:
```bash
echo "$NODUS_API_KEY" | nodus login --with-token
```
Prefer a [service account token](/docs/guides/access/service-accounts/) for CI, so it keeps working when people leave.
# Members and roles
> Invite people to your org, choose their role, and remove members safely.
## Roles
[Section titled “Roles”](#roles)
| Role | Can |
| ---------- | ------------------------------------------------------------------------------------------------- |
| **Owner** | Everything, including granting or removing the Owner role |
| **Admin** | Everything except the Owner role: members, invites, projects, API keys, service accounts, billing |
| **Member** | Create and manage work: jobs, sandboxes, volumes, secrets, their own API keys |
| **Viewer** | Read everything, change nothing |
`nodus auth can-i create jobs` asks the server whether you hold a permission.
## Invite someone
[Section titled “Invite someone”](#invite-someone)
```bash
nodus create invite --email ada@example.com --role Member
```
The invite link works for 72 hours and only once. The invitee signs in with that email address (it must be verified) and accepts it in the console, where pending invites also appear as a banner. You can invite up to your seat limit, 10 members by default; pending invites count as seats. An invite nobody accepts stays under **Team › Invites** after it expires, marked Expired and holding no seat, so you can resend it. Once your org has bought credits, an Owner or Admin can change the seat limit under **Team › Members › Change**, up to 1,000. When an invite is accepted, or someone is made an Admin or Owner, every Owner and Admin gets an email.
```bash
nodus get invites
nodus request resend invite/inv-01j9abc # a fresh link, at most 3 times, 10 minutes apart
nodus delete invite inv-01j9abc
```
You can invite someone with a role up to your own: Admins invite Admins, Members and Viewers; only Owners invite Owners.
## Let your company join without invites
[Section titled “Let your company join without invites”](#let-your-company-join-without-invites)
An Owner or Admin adds your company’s email domain under **Team › Company domains**, then adds the TXT record shown there at your DNS provider and selects **Verify**. After that, anyone who signs up with a verified email at that domain is offered your org during sign-up and joins it as a Member in one click. Seats still apply, public mailbox domains such as `gmail.com` are refused, and a domain belongs to one org at a time. Remove the domain to stop new joins; people who already joined stay members. Someone you remove from the org can only come back through an invite.
## Change a role
[Section titled “Change a role”](#change-a-role)
```bash
nodus get members
nodus edit member usr-01j9abc --role Admin
```
An org always keeps at least one Owner: demoting or removing the last Owner is refused with `409 Conflict`.
## Remove a member
[Section titled “Remove a member”](#remove-a-member)
```bash
nodus delete member usr-01j9abc
nodus delete member usr- # leave an org yourself
```
Removing a member revokes their API keys and third-party app access in that org at once. Keys bound to a [service account](/docs/guides/access/service-accounts/) keep working, so CI does not break when the person who set it up leaves.
# Account security
> Sign in and control access with organization roles and scoped credentials.
Sign in with your email and password or an enabled Google or GitHub account. You can then use the resources and settings your organization role allows. Nodus does not require QR scanning, an authenticator app or a separate verification code after sign-in. Email verification for new accounts and password recovery still apply.
CLI and device login use your existing signed-in session. API keys and service account tokens keep their own scopes. Review them under [API keys](/docs/guides/access/api-keys-and-scopes/).
# Projects
> Group work inside an org with projects, and restrict members and keys to some of them.
A **project** is a namespace inside your org. Jobs, Sandboxes, Volumes, Secrets and the other resources you create live in one project. Every org has `default`, which cannot be deleted.
```bash
nodus create project research --display-name "Research"
nodus get projects
nodus run -p research -- python train.py
```
In the console, **Team › Projects** lists each project with how many members can use it, its spend over the last 30 days and its spend limit, and has **New project**. Give each team its own project: its runs, storage and keys stay apart, and its spend shows on its own line.
A project name is lowercase letters, digits and dashes. `nodus`, `team`, `settings`, `billing`, `integrations`, `infrastructure` and `support` are reserved.
Creating, changing and deleting projects needs the `projects:write` scope, which Admins and Owners hold. Deleting a project deletes everything in it: running work is cancelled first. Usage and billing records are kept.
## Restrict access to projects
[Section titled “Restrict access to projects”](#restrict-access-to-projects)
Members and API keys can be limited to some projects:
```bash
nodus edit member usr-01j9abc --projects research
nodus create apikey ci --projects research --scopes jobs:write
```
A member or key with a project list sees and changes only those projects. An empty list means every project.
## Cap a team’s spend
[Section titled “Cap a team’s spend”](#cap-a-teams-spend)
A budget scoped to a project caps that team. From **Team › Projects**, open a project’s menu and choose **Set a spend limit**, or:
```bash
nodus create budget vision-monthly --limit 500 --scope-project vision-team
```
Budgets can also cap one person, or one person inside one project. When several budgets cover the same work, the tightest one applies.
# Quotas
> See the limits on your org, what counts against them, and how to raise them.
Quotas keep one mistake from running away with your credits. See yours with:
```bash
nodus get quota default -o yaml
```
| Quota | Default before your first purchase | After a purchase |
| --------------------------------------------------------------------- | ---------------------------------- | ---------------- |
| `members` (seats, including pending invites) | 10 | 10 |
| `apiKeys` (live keys, including CLI logins and ServiceAccount tokens) | 50 | 200 |
| `liveSandboxes` (stopped ones count until their delete completes) | 10 | 100 |
| `storageGiB` | 50 | 1024 |
| `egressGiBPerDay` | 20 | 100 |
| `nodusNodes` | 1 | 10 |
| `gangNodes` (multi-node jobs) | 0 | 8 |
| `inferenceRPM` / `inferenceTPM` | 60 / 100,000 | 600 / 1,000,000 |
| `assistantUSDPerDay` | $1 | $10 |
Going over a count quota refuses the create with `429 QuotaExceeded`, naming the quota. Nothing already running is touched. Daily quotas (`egressGiBPerDay`, `assistantUSDPerDay`) reset at 00:00 UTC. Expired and revoked API keys do not count against `apiKeys`.
Multi-node jobs are in Beta: your org needs the distributed training beta, and `gangNodes` counts the nodes of running multi-node jobs. A job that would go over it waits in the queue instead of failing.
Owners and Admins of an org that has bought credits set their own seats, up to 1,000, under **Team › Members › Change**, or with `PATCH /apis/nodus.dev/v1/quotas/default` and `{"hard": {"members": 25}}`. Seats never go below the members and pending invites you have. To raise any other org limit, or seats past 1,000, contact support. To cap what one project or one person spends, use a [Budget](/docs/guides/billing/#cap-spending-with-a-budget).
# Service accounts
> Give CI and code running on Nodus an identity that belongs to the project, not to a person.
A **service account** is an identity for machines. Every project has one named `default`, used by code running inside Nodus for its own calls. Create more for CI and automation:
```bash
nodus create serviceaccount ci -p research --scopes jobs:write,volumes:read
```
Managing service accounts needs `serviceaccounts:write` (Admins and Owners).
## Tokens for CI
[Section titled “Tokens for CI”](#tokens-for-ci)
```bash
nodus create token sa/ci -p research --duration 720h
```
This prints a key bound to the service account, once. It is valid for 90 days by default and at most one year. Its permissions are the service account’s scopes, narrowed further by any `--scope` you pass.
Service account keys do not belong to a person: they keep working when the member who created them leaves the org, and they stop working when the service account is deleted. Use them for the GitHub Action and any shared automation.
## Inside Nodus
[Section titled “Inside Nodus”](#inside-nodus)
Code running in a Job, Sandbox or Agent reaches the API through `/run/nodus/api.sock` with a short-lived token of its service account. The token is never in the environment or on disk, and it stops working the moment its run is stopped or moved.
# Sign up and create your org
> Create an account, create or join an organization, and switch between the orgs you belong to.
Everything you run on Nodus belongs to an **organization** (org). Credits, projects, members and API keys are all per org.
## Sign up
[Section titled “Sign up”](#sign-up)
Sign up at the console with email, Google or GitHub. Nothing is created until you choose:
* If someone invited you, the console shows the invite first. Accept it to join their org.
* Otherwise the console asks you to **Create your org**. The name is filled in for you, and you can change it.
The first org you create with a verified email address receives the **$30 starter credit**, valid for 30 days. Joining an org by invite never adds credit, so accepting an invite does not use up your own starter credit.
## Start from the terminal
[Section titled “Start from the terminal”](#start-from-the-terminal)
```bash
nodus login
```
This opens the console in your browser. Sign in, pick the orgs this computer should use, and the CLI stores one key per org in your OS keychain. If you have no org and no pending invite, `nodus login` creates your org for you, with the starter credit. On a machine without a browser, use [device login](/docs/guides/access/device-login/).
## Create another org
[Section titled “Create another org”](#create-another-org)
```bash
nodus create org acme --display-name "Acme Research"
```
You become its Owner. Every org starts with a project named `default`.
## Switch orgs
[Section titled “Switch orgs”](#switch-orgs)
The console has an org switcher. On the CLI, each org is a context:
```bash
nodus config get-contexts
nodus config use-context acme
nodus get jobs --org acme # one command in another org
```
API calls choose an org with the `Nodus-Org` header (the org name or id). Without it, a signed-in user acts in the org they used last.
## Org settings and activity
[Section titled “Org settings and activity”](#org-settings-and-activity)
Admins can rename the org; its name (the slug in console URLs) stays the same:
```bash
nodus patch org acme --patch '{"displayName": "Acme Labs"}'
nodus get org acme -o yaml # role, seats used and your org's enabled features
```
Admins and Owners see the org’s activity under **Settings → Activity**: sign-ins, key, member and invite changes, each with who did it and when, for the last day up to the last year, and **Export CSV** downloads what you see. The API serves it at `GET /apis/nodus.dev/v1/auditevents`, oldest first: `?action=auth.` selects sign-ins only, `since` sets the start, and a full page carries `metadata.continue` to pass as `continue` for the next one.
# SSH keys
> Add the public keys you use to open a shell in a Workspace.
An **SSH key** belongs to you, not to an org. Add it once and it works in the Workspaces of every org you belong to. Console › Account › SSH keys lists yours whatever org is selected.
```bash
nodus create sshkey laptop --from-file ~/.ssh/id_ed25519.pub
nodus get sshkeys
```
Without a key yet, let the CLI make one. `--generate` writes a new Ed25519 key pair to `~/.ssh/id_ed25519` (the public key next to it, with `.pub`) and adds the public key. The private key file is readable by you alone, and `nodus` never sends it anywhere. It refuses to overwrite a file that is there: use `--key-file PATH` to write elsewhere. Add a passphrase afterwards with `ssh-keygen -p -f ~/.ssh/id_ed25519`.
```bash
nodus create sshkey --generate
```
Leave the name out and the key is called `-`. `--from-file -` reads the key from standard input, and a file that holds a private key is refused.
Or over the API:
```bash
curl -X POST "$NODUS_API_URL/apis/nodus.dev/v1/sshkeys" \
-H "Authorization: Bearer $NODUS_API_KEY" -H "Content-Type: application/json" \
-d "{\"name\": \"laptop\", \"publicKey\": \"$(cat ~/.ssh/id_ed25519.pub)\"}"
```
Nodus accepts `ssh-ed25519`, `ecdsa-sha2-*` and `ssh-rsa` keys of at least 3072 bits. It stores the key without its comment and shows its `SHA256:` fingerprint, which is what `ssh-keygen -lf ~/.ssh/id_ed25519.pub` prints. Each name and each key can be added once, and you can hold up to 20 keys.
## Who can connect
[Section titled “Who can connect”](#who-can-connect)
`nodus ssh workspace/` authenticates with your keys when you are an Owner, Admin or Member of the Workspace’s org and your membership covers its project. Viewers cannot open a shell. When you leave an org, your keys stop working in its Workspaces, and they keep working everywhere else.
## Removing a key
[Section titled “Removing a key”](#removing-a-key)
```bash
nodus delete sshkey laptop
```
Deleting a key ends the live sessions opened with it, and the next login with it is refused.
# Agents
> Run an agent on Claude in its own sandbox, send it follow-up prompts, and pay for the model from your credits.
An **Agent** is a definition: a system prompt, the Claude access it runs on and a cost cap. An **AgentRun** is one conversation with it. Each run works in its own sandbox, where the model runs shell commands to check its work, and it survives restarts of Nodus by replaying recorded results. An interrupted command or model call using your own key parks with `NeedsResolution` when its outcome is uncertain; inspect its effects before starting replacement work.
Agents, parallel AgentGroups and Environment evaluation batches are in Beta. Each run is one agent conversation.
## Sign in and create an agent
[Section titled “Sign in and create an agent”](#sign-in-and-create-an-agent)
agent.yaml
```yaml
apiVersion: nodus.dev/v1
kind: Agent
metadata:
name: hello-agent
spec:
system: |
You are a careful assistant. Run shell commands in your sandbox to check your work, and answer in one short
paragraph.
# Nodus's Claude: each prompt is routed to Haiku, Sonnet, Opus or Fable, and billed from your credits.
model:
access: Nodus
perRunMaxCostUSD: "0.50"
```
```bash
nodus login
nodus apply -f agent.yaml
```
The agent uses Nodus’s Claude. Every prompt goes to a small router that picks the right Claude model for it (Haiku for simple requests up to Fable for demanding ones). That model answers the whole prompt, including its commands, so the model never changes in the middle of a turn. If the router is unavailable, Sonnet answers. Restrict the choice with `spec.model.families`, for example `[haiku, sonnet]`.
## Start a run
[Section titled “Start a run”](#start-a-run)
run.yaml
```yaml
apiVersion: nodus.dev/v1
kind: AgentRun
metadata:
name: hello-agent-run
spec:
agent: hello-agent
input: "Use the shell to print the Python version, then tell me which one it is."
```
```bash
nodus apply -f run.yaml
nodus wait agentrun/hello-agent-run --for=jsonpath='{.status.phase}'=Succeeded --timeout=8m
nodus get agentrun/hello-agent-run -o yaml
```
Or without a file: `nodus create agentrun triage --agent hello-agent --prompt "Summarize the logs"`. To start from Nodus’s ready-made assistant, name the template `claude-assistant`.
`status.turns` lists each prompt with the model that served it, how it was chosen (`jev` or `fallback`) and its cost. `status.answer.preview` holds the first 4 KiB of the final answer; `nodus agentrun answer` prints all of it.
```bash
nodus agentrun answer hello-agent-run
nodus agentrun steps hello-agent-run
```
`steps` lists what the run did in order: creating its sandbox, each routing and model call, each command. A step that shows `Completed` is never repeated, even after a restart.
The console’s Agents page shows the same: each run with the model of its latest turn and its cost, and on a run every turn with the model that served it, how that model was chosen and what the turn cost, then the answer and the recorded steps.
## From Python
[Section titled “From Python”](#from-python)
```python
import nodus
agent = nodus.ClaudeAgent("helper", system="You are careful.", families=["haiku", "sonnet"])
print(agent.remote("Use the shell to print the Python version")) # runs to the end, returns the answer
run = agent.submit("Summarize the logs", keep_alive=True) # a handle: run.answer(), run.steps(), run.cancel()
run.send("message", "Now the staging logs")
```
`ClaudeAgent` creates the agent on first use. Pass `api_key_secret="anthropic-key"` to use your own Anthropic key, or `ClaudeAgent.from_name("claude-assistant", project="nodus")` to run the ready-made template. To run many at once from Python, see [Run agents in parallel](/docs/guides/python/agents/#run-agents-in-parallel).
## Keep a run open for follow-ups
[Section titled “Keep a run open for follow-ups”](#keep-a-run-open-for-follow-ups)
Set `spec.keepAlive: true` and the run waits for messages after each turn instead of finishing. While it waits it uses no model and keeps its sandbox, which can still accrue compute charges. Set `spec.deadline` to bound the run’s lifetime. The deadline also applies while waiting for a message: the run fails with `DeadlineExceeded` and its sandbox is deleted without another model call.
```bash
nodus apply -f run.yaml # with keepAlive: true
nodus agentrun send hello-agent-run "Now do the same for the staging logs" --key staging-1
```
Sending the same `--key` again delivers the message once. A run that is still working reads the message after its current turn; a finished run refuses it with `AgentRunFinished`. Cancel a run with `nodus cancel agentrun/NAME`; its sandbox is deleted.
The same calls are REST: `POST …/agentruns/NAME/messages` with `{"payload": "…", "messageKey": "…"}`, and `GET` on `…/steps` and `…/answer`.
## Run agents in parallel
[Section titled “Run agents in parallel”](#run-agents-in-parallel)
An **AgentGroup** runs many runs of one agent at once. You choose how many run at the same time and how much the whole group may spend. The other runs wait their turn, and one command cancels or deletes all of them. Create the agent and the group, then one run for each task:
agent.yaml
```yaml
apiVersion: nodus.dev/v1
kind: Agent
metadata:
name: parallel-agents-worker
spec:
system: |
Follow the task's requested answer format; otherwise answer in one short sentence.
# Nodus's Claude, limited to Haiku so each run stays far below its cap.
model:
access: Nodus
families: [haiku]
maxTokens: 1024
perRunMaxCostUSD: "0.10"
```
group.yaml
```yaml
apiVersion: nodus.dev/v1
kind: AgentGroup
metadata:
name: parallel-agents
spec:
agent: parallel-agents-worker
limits:
maxActive: 2 # at most two runs go at once
maxCostUSD: "0.60" # a run starts only while its own cap still fits under what is left
```
runs.yaml
```yaml
apiVersion: nodus.dev/v1
kind: AgentRun
metadata:
name: parallel-agents-a
spec:
agent: parallel-agents-worker
group: parallel-agents
taskKey: a
input: "What is the capital of France?"
---
apiVersion: nodus.dev/v1
kind: AgentRun
metadata:
name: parallel-agents-b
spec:
agent: parallel-agents-worker
group: parallel-agents
taskKey: b
input: "What is the capital of Japan?"
---
apiVersion: nodus.dev/v1
kind: AgentRun
metadata:
name: parallel-agents-c
spec:
agent: parallel-agents-worker
group: parallel-agents
taskKey: c
dependsOn: [a, b] # c waits until a and b have succeeded
input: "Say in one sentence that both questions have been answered."
```
```bash
nodus apply -f agent.yaml
nodus apply -f group.yaml
nodus apply -f runs.yaml
nodus patch ag/parallel-agents --patch '{"spec":{"sealed":true}}'
```
Or without a file: `nodus create agentgroup parallel-agents --agent parallel-agents-worker --max-active 2 --max-cost 0.60`.
Each run is a task of the group: `spec.group` names the group and `spec.taskKey` names the task (a DNS label that is unique in the group). The run is named `-`; leave `metadata.name` out or set exactly that. The run uses the group’s agent, so `spec.agent` must match it. The last command seals the group, which says that no more runs are coming (see below).
### Watch progress and read the answers
[Section titled “Watch progress and read the answers”](#watch-progress-and-read-the-answers)
```bash
nodus wait ag/parallel-agents --for=jsonpath='{.status.phase}'=Succeeded --timeout=12m
nodus get ag
nodus get ar --field-selector spec.group=parallel-agents
nodus agentrun answer parallel-agents-c
```
`nodus get ag` lists each group with its phase, `RUNS` (finished successfully out of all) and `ACTIVE` runs, and its cost. `status.counts` splits the runs into `queued`, `active`, `waiting` (for a message or for funds), `succeeded`, `failed` and `cancelled`. `status.blockedReasons` says why queued runs have not started and how many wait for each reason: `DependencyWait`, `MaxActiveReached` or `MaxCostReached`. A queued run’s own `status.reason` is `DependencyWait`, or `GroupWait` while it waits for a free slot or for room under the cost cap. The runs you list are ordinary runs, so `status.answer`, `nodus agentrun answer` and `nodus agentrun steps` work on each of them.
A group is `Running` until it is sealed and every run has finished. It then becomes `Succeeded`, or `Failed` with reason `RunsFailed` when a run failed or was cancelled because a dependency failed. A sealed group with no runs succeeds.
### How many run at once
[Section titled “How many run at once”](#how-many-run-at-once)
`spec.limits.maxActive` is how many runs the group lets run at the same time (default 10, 1 to 1000). Change it on a live group:
```bash
nodus patch ag/parallel-agents --patch '{"spec":{"limits":{"maxActive":4}}}'
```
`spec.limits.maxPending` is how many unfinished runs the group takes in (default 10000); a run beyond it is refused.
### Dependencies
[Section titled “Dependencies”](#dependencies)
`spec.dependsOn` lists the task keys a run waits for. It starts only after every one of them has succeeded, and until then it is queued with reason `DependencyWait`. A dependency must already be in the group, so create the runs it depends on first (earlier in the same file works). When a dependency fails, the run is cancelled with reason `DependencyFailed`; when a dependency was cancelled, the run is cancelled with reason `DependencyCancelled`.
### The group cost cap
[Section titled “The group cost cap”](#the-group-cost-cap)
`spec.maxCostUSD` is one cap for the whole group. A run starts only when its own cap (the agent’s `perRunMaxCostUSD`) still fits under what is left: the cap minus what the group has spent and minus the caps of the runs that have started and not finished. The members together can therefore never spend more than the cap, and runs that do not fit yet wait with `MaxCostReached`. In the example, the $0.60 cap has room for the caps of six runs at $0.10 each before any run has spent anything. Raise the cap on a live group with `nodus patch`; it cannot be lowered. If the cap is smaller than one run’s cap, no run starts: raise it or lower the agent’s `perRunMaxCostUSD`. `status.cost` shows `totalUSD` and `limitUSD`. The cap covers the runs’ Claude usage (model and routing calls), the spend each run’s `perRunMaxCostUSD` limits; the sandbox each run works in is billed on its own and is not counted against either cap.
### Seal, cancel and delete
[Section titled “Seal, cancel and delete”](#seal-cancel-and-delete)
A group takes new runs until you seal it with `spec.sealed: true`. Sealing cannot be undone, and a sealed group refuses new runs with a conflict. A group finishes only once it is sealed, so seal it after you create the last run.
```bash
nodus cancel ag/parallel-agents
nodus delete ag/parallel-agents
```
`nodus cancel` cancels every run that has not finished and leaves the finished runs, with their answers, as they are. `status.phase` goes through `Cancelling` to `Cancelled`. `nodus delete` removes the group together with all of its runs.
## Evaluate an agent
[Section titled “Evaluate an agent”](#evaluate-an-agent)
Create an evaluation group with an Environment at an explicit version. Nodus generates the fixed task batch, runs it through ordinary AgentRuns, and grades completed answers using the Environment’s isolated grader. There is no manual submit or seal step. Set a project Budget to limit total spend: the group cap covers model usage, while agent and grading Sandbox compute is billed separately.
After deploying the example’s `parallel-agents-worker` Agent above, run this Python example. It creates two held-out tasks, waits for a terminal result, prints per-case outcomes and deletes the evaluation group afterward.
evaluation.py
```python
"""Evaluate the deployed example agent on a fixed, versioned task batch."""
import nodus
def main() -> None:
# Model caps exclude Sandbox compute; set a project Budget before running this example.
evaluation = nodus.AgentGroup.create(
"parallel-agents-eval",
agent="parallel-agents-worker",
max_active=2,
max_cost="0.20",
evaluation={
"environment": "nodus/arithmetic-v2@2.0.0",
"split": "test",
"tasks": 2,
"repetitions": 1,
"seed": 42,
"timeout": "10m",
},
)
try:
status = evaluation.wait(timeout=660)
results = evaluation.results()
print(
f"{status.phase}: cases={len(results)}, pending={status.evaluation.pending}"
)
print("pass rate:", status.evaluation.get("passRate"))
print("mean reward:", status.evaluation.get("meanReward"))
for result in results:
print(result.task_id, result.state, result.get("verdict"))
finally:
evaluation.delete()
if __name__ == "__main__":
main()
```
The batch is at most 100 cases including repetitions. Every member pins the same Agent revision; results record the Environment version, resolved image digest, split and seed. Only task prompts reach the agent, never hidden answers. The timeout defaults to 30 minutes, may be 1 minute to 24 hours, and includes manifest preparation, queueing and funding waits from group creation. Each member receives that absolute deadline. Grading also stops at the deadline; completed answers that could not be scored are shown as evaluation timeouts, excluded from pass rate and mean reward.
Open the group in the console to see scored cases, pass rate, mean reward, pending cases and execution/grading failures separately. The Results view links each case to its AgentRun and shows grader evidence. A failed run, missing answer or grader failure is unscored and does not become a zero reward. A wrong or invalid answer is a scored result. Per-case records share the member journal’s 30-day payload retention; deleting the group removes its runs and records. Canceling or finishing the group cleans up its grading pool.
## Use your own Anthropic key
[Section titled “Use your own Anthropic key”](#use-your-own-anthropic-key)
Store the key in a Secret under the key `ANTHROPIC_API_KEY` and point the agent at it. Nodus then charges nothing for the model, and your Anthropic account pays for it.
```yaml
spec:
model:
access: BYOK
name: claude-sonnet-5-5
apiKeySecret: anthropic-key
```
## What a run costs
[Section titled “What a run costs”](#what-a-run-costs)
On Nodus’s Claude each model call, and the routing call before each turn, is billed from your credits at cost divided by 0.875. `spec.perRunMaxCostUSD` (default $1.00) caps what a run spends on the model. Before each call the run checks that the call, at its largest (the whole conversation as input and `spec.maxTokens` of output), still fits under the cap. When it does not, the run waits (`reason: MaxCostReached`) and makes no more calls. The cap is copied from the agent when the run is created, so changing the agent does not change a run that has started: to continue the work, cancel the run and start a new one. If your credits run out, the run waits instead and continues when you add credits. A run on your own Anthropic key is not charged by Nodus and has no cap to reach. Sandbox compute is billed as for any sandbox. `status.cost` shows the model spend so far. Runs in a group also share the group’s cap (see [Run agents in parallel](#run-agents-in-parallel)).
## Limits
[Section titled “Limits”](#limits)
* One run works one prompt at a time, with at most `spec.maxToolCalls` commands per turn (default 50).
* The input and each message are at most 256 KiB. A run’s recorded payloads are kept 30 days after it ends.
* A run created for a stopped agent (`spec.state: Stopped`) is refused with `AgentStopped`.
* A group takes at most 10000 unfinished runs. Create runs in batches of up to 100 by posting an `AgentRunList` to `…/agentruns` with an `Idempotency-Key` header; a batch is all-or-nothing, and its runs are created in list order, so list each run after the runs it depends on.
See [Durable execution](/docs/concepts/durable-execution/) for what survives a restart.
# Ask Nodus
> Ask the console assistant why a Job failed, what it cost, or to draft a manifest. It cites the docs, and changes only after you confirm.
**Ask Nodus** is the assistant in the console. It answers questions about your org: why a Job restarted, what a Sandbox cost this week, which flags a manifest needs. It uses the same tools as the [MCP server](/docs/guides/mcp/), with your own role and scopes, so it can see and do only what you can.
## Ask a question
[Section titled “Ask a question”](#ask-a-question)
Open Ask Nodus in the console and type. Some things to try:
* “Why did job/train fail?”
* “What did my Sandboxes cost this month?”
* “How do I make train survive being preempted?”
* “Write a manifest that runs `python train.py` on an L4 with a $5 cap.”
When an answer relies on the docs, it links the pages, and the links are the pages the assistant actually read.
## Changes need your confirmation
[Section titled “Changes need your confirmation”](#changes-need-your-confirmation)
The assistant can propose a change, but it cannot make one. When an answer would create, delete, suspend or otherwise change something, the turn stops and shows exactly what will run. For a new object that includes the server’s dry-run: the object as it would be stored, its estimated cost and anything that blocks it.
Choose **Confirm** to run it, or **Cancel**. Confirming runs the held change once, with your credential, so the API checks your permissions again. Asking a new question cancels a change you did not confirm. The assistant proposes at most one change per answer.
## Drafts
[Section titled “Drafts”](#drafts)
When you ask for a manifest, the assistant writes one and checks it with the same dry-run that `nodus apply --dry-run=server` uses. A draft is shown to you only after that check passes. If the check fails, the assistant sees the error and fixes the manifest first. Each draft comes with the equivalent CLI command and Python, and its estimate. A draft is not created until you apply it, or ask the assistant to and confirm.
## Checkpoint suggestions
[Section titled “Checkpoint suggestions”](#checkpoint-suggestions)
For a Job that would start over after a preemption, the assistant can suggest a manifest that checkpoints, with that manifest’s dry-run. Recovery settings cannot change on a running Job, so the suggestion is a new manifest. Your program still has to write its state to the declared checkpoint paths.
## Your preferences and followed objects
[Section titled “Your preferences and followed objects”](#your-preferences-and-followed-objects)
You can give the assistant a short note about how you like answers, such as “short answers, show the CLI command”. It reads the note as your preference, and the note cannot change the assistant’s rules. You can also follow up to 8 objects so the assistant knows what you are working on.
Both are yours alone in each org. In the API:
```console
$ curl -s "$NODUS_API_URL/assistant/v1/profile" -H "Authorization: Bearer $NODUS_API_KEY"
$ curl -s -X PUT "$NODUS_API_URL/assistant/v1/profile" -H "Authorization: Bearer $NODUS_API_KEY" \
-H "Content-Type: application/json" -d '{"note": "Short answers.", "revision": 0}'
```
The note is at most 8 KiB. Send back the `revision` you read, and the update fails with a conflict if someone else changed the note in between. `/assistant/v1/watches` takes `{"watches": ["", …]}`.
## Chats and privacy
[Section titled “Chats and privacy”](#chats-and-privacy)
Your chats are stored per org, for you only. A chat holds up to 200 questions. Each answer is stored with a checksum, and deleting a chat deletes its answers. A tool the assistant ran is recorded by name with checksums of its arguments and its result, never their content, so file contents and command output are not kept. Chats are kept at most 400 days.
Nodus chooses a Claude model for each question and uses it for the whole answer. Nodus pays for the model calls; they do not use your credits.
## Limits
[Section titled “Limits”](#limits)
| Limit | Value |
| ---------------------- | --------------------------------------------------------------------------- |
| Questions | 1 per second per user, with bursts of up to 20 |
| Steps for one question | 12 model rounds and 32 tool calls |
| Time for one question | 180 seconds |
| Drafts | 1 draft request every 5 seconds per user |
| Daily allowance | Each org has a daily allowance for the assistant, which resets at 00:00 UTC |
When a limit is reached the assistant says so and what happens next, instead of failing. Your Jobs, the docs and the CLI are not affected.
Note
The assistant reads logs and files as data, never as instructions. Text inside a log cannot make it run anything, and anything it proposes still needs your confirmation.
# Usage and billing
> Add credits, redeem a code, cap spending with Budgets, and see exactly what every run cost.
Nodus is prepaid: you add credits, and every run reserves funds before it starts. New orgs get [starter credit](/docs/guides/billing/credits/) to try things out. This guide covers the everyday tasks; the [billing concept page](/docs/concepts/billing/) explains holds, captures and limits.
## Check your balance
[Section titled “Check your balance”](#check-your-balance)
```console
$ nodus billing
Available $17.42 purchased $15.80 · credits $3.52 (starter, expires Oct 31)
Reserved $1.90 job/train-a, sandbox/sb-3, function/embed
This month $232.58 budget research-monthly 42 % of $500.00
Auto-recharge on add $50.00 below $10.00 · visa •••• 4242
```
In the console, open **Usage & billing**. The header shows your available balance everywhere, in amber when it is low.
## Add credits
[Section titled “Add credits”](#add-credits)
```console
$ nodus billing top-up 20
Opening https://checkout.stripe.com/c/pay/cs_live_... (expires in 60 min)
Waiting for payment... added $20.00. Available $37.42.
Receipt: https://invoice.stripe.com/i/...
```
Top-ups are between $5.00 and $1,000.00 in whole cents, up to $5,000 per org per day. Payment happens on a Stripe-hosted page; Nodus never sees your card number. In the console, **Add credits** offers $10, $20, $50, $100 or a custom amount and brings you back to Billing when the payment completes. A top-up settles any unpaid charges first. Every top-up has a receipt and an invoice PDF under **Receipts & invoices** and in `nodus billing receipts`.
Topping up needs the `billing:write` permission, which org Owners and Admins have.
## Auto-recharge
[Section titled “Auto-recharge”](#auto-recharge)
Auto-recharge adds credit when your available balance drops below a threshold, so long runs never stop for money:
```console
$ nodus billing auto-recharge --threshold 10 --amount 50
```
The first time, this opens a card setup page. Each recharge is a Stripe invoice with its own receipt. At most 5 recharges or $5,000 run per day, and three failed charges in a row turn auto-recharge off and email you. Turn it off with `nodus billing auto-recharge --off`.
**Save** updates the amounts and warning level without turning auto-recharge on or off. Use **Turn on** or **Turn off** to change that setting. If another admin or failed payments change it while you edit, the console refreshes the settings and asks you to review them before saving again.
## Redeem a promo code
[Section titled “Redeem a promo code”](#redeem-a-promo-code)
```console
$ nodus billing redeem LAUNCH25
Redeemed ****CH25: $25.00 of credit, expires 2026-11-30.
```
See [Promo codes](/docs/guides/billing/promo-codes/) for limits and errors.
## See what you spent
[Section titled “See what you spent”](#see-what-you-spent)
```console
$ nodus get usage --group-by project,label:owner --since 30d
PROJECT OWNER AMOUNT
research ml $212.41
default - $20.17
$ nodus get usage --group-by project,meter --since 30d -o csv > usage.csv
$ nodus get transactions --since 7d
```
Group usage by `project`, `kind`, `label:`, `meter`, `day`, `segment` or `rank`. `segment` splits machine time into `Boot`, `Restore`, `Running` and `Teardown`, and `rank` itemizes the members of a multi-node run. The console **Usage** tab shows the same data as a daily chart and a table, and every table exports CSV.
Transactions list every change to your balance, with the balance after it: top-ups, grants, captures, storage, egress, refunds and adjustments. [What you pay for](/docs/guides/billing/what-you-pay-for/) lists every kind of time and who pays for it.
## Cap spending with a Budget
[Section titled “Cap spending with a Budget”](#cap-spending-with-a-budget)
A Budget is an enforced limit over the org, a project or a label selector, per month or in total. When a `Block` Budget is exhausted, work in its scope stops gracefully and new work is refused until the next period:
budget.yaml
```yaml
apiVersion: nodus.dev/v1
kind: Budget
metadata:
name: examples-billing-monthly
spec:
limitUSD: "25.00"
period: Monthly
scope:
project: examples
action: Block
thresholds: [50, 80, 100]
```
```console
$ nodus apply -f budget.yaml
$ nodus get budgets
```
You get an email at each threshold. The console **Budgets** tab creates and edits Budgets and previews which objects a scope matches now.
To cap one person, scope a Budget to them by email. It counts everything they start, with any of their API keys, including inference requests:
```console
$ nodus create budget ada-monthly --limit 300 --period Monthly --scope-member ada@example.com
```
## When work is refused for money
[Section titled “When work is refused for money”](#when-work-is-refused-for-money)
A create that cannot be funded fails with `402` and the exact amounts, and the console shows an **Add credits** action:
| Error | What to do |
| --------------------- | --------------------------------------------------------------- |
| `InsufficientCredits` | Add credits, or lower `spec.maxCostUSD` |
| `BudgetExceeded` | Raise the Budget’s `spec.limitUSD`, or wait for its next period |
| `ArrearsOutstanding` | Add credits; the top-up settles the unpaid charges first |
| `PaymentDisputed` | Contact support; new work waits until the dispute closes |
## Payment methods and billing details
[Section titled “Payment methods and billing details”](#payment-methods-and-billing-details)
`nodus billing portal` opens the Stripe Customer Portal, where you update cards, your billing email, address and tax ID, and download past invoices.
## Learn more
[Section titled “Learn more”](#learn-more)
* [Credits](/docs/guides/billing/credits/): the starter credit, grants and expiry
* [Promo codes](/docs/guides/billing/promo-codes/)
* [Refunds](/docs/guides/billing/refunds/)
* [What you pay for](/docs/guides/billing/what-you-pay-for/)
# Add a card before running work
> Save a payment method without buying credits.
Open **Billing → Payment method → Add a card**, or run:
```sh
nodus billing portal --setup
```
Stripe saves your card without charging it or enabling auto-recharge. Your $30 signup credit keeps its original 30-day expiration. New work needs a verified card even when promotional credit covers its entire cost. An organization billing administrator must add the card. Other members can view its status.
After returning from Stripe, wait for Billing to show that the card is ready. Returning to the page alone does not complete verification. If you canceled setup, choose **Add a card** again. CLI and SDK requests receive `PaymentMethodRequired` until verification completes. `PaymentVerificationUnavailable` means to retry shortly.
Removing or letting the card expire blocks new work. Existing running work continues within its available credits and budgets. A restart or new allocation checks the payment method again. Adding a card does not let work exceed your prepaid balance.
# Auto-recharge
> Top up automatically from a saved card when your balance falls below a threshold, so running work never stops for money.
Auto-recharge buys credits for you when your available balance falls below a threshold you choose. Each recharge is charged to a saved card and emailed to you as a receipt, exactly like a top-up you buy yourself. It is off until you turn it on.
## Turn it on
[Section titled “Turn it on”](#turn-it-on)
```bash
nodus billing auto-recharge --threshold 10 --amount 50
```
This recharges $50 whenever your available balance drops below $10. If no card is on file yet, the command first opens a Stripe page to save one, then turns auto-recharge on. You can also save a card while buying credits with `nodus billing top-up 20 --save-card`, and the console billing page has an auto-recharge card that does both.
The same settings live on your org’s `BillingAccount`:
```yaml
apiVersion: nodus.dev/v1
kind: BillingAccount
metadata:
name: default
spec:
autoRecharge:
enabled: true
thresholdUSD: "10.00" # at least 5.00
amountUSD: "50.00" # 5.00 to 1000.00
```
| Setting | Default | Allowed |
| -------------- | --------- | ------------------ |
| `thresholdUSD` | `"10.00"` | $5.00 or more |
| `amountUSD` | `"20.00"` | $5.00 to $1,000.00 |
Turning it on without a saved card is refused; check `PaymentMethodPresent` in `nodus billing`, or save a card through the billing portal (`nodus billing portal`).
## When it recharges
[Section titled “When it recharges”](#when-it-recharges)
Nodus checks your balance each time it renews the hold of running work, settles inference or storage, or refuses a new hold. It recharges when all of these are true:
* your available balance is below `thresholdUSD`;
* a card is on file;
* no other auto-recharge is still being paid;
* the org has had fewer than 5 auto-recharges, and less than $5,000 of them, today (UTC);
* the last 3 auto-recharges did not all fail.
A recharge usually lands in seconds. Work whose funding is running short keeps running while it is paid, so a successful recharge within that window means nothing stops.
## See what it did
[Section titled “See what it did”](#see-what-it-did)
Each recharge is a `TopUp` with `spec.origin: AutoRecharge`:
```bash
nodus get topups -o wide
nodus billing receipts
```
`nodus billing` shows the auto-recharge state from `status.autoRecharge`: `consecutiveFailures`, `lastAttemptTime`, `lastResult` and, when it has been switched off, `disabledReason`.
## When a recharge fails
[Section titled “When a recharge fails”](#when-a-recharge-fails)
If the card is declined, the recharge `TopUp` ends `Failed` (`PaymentFailed`), nothing is credited, and the billing email and the org’s owners and admins get an email. Auto-recharge tries again the next time the balance is checked.
After 3 failures in a row, Nodus turns auto-recharge off, sets `status.autoRecharge.disabledReason` to `ConsecutiveFailures`, shows a banner in the console and emails the same people. To turn it back on:
1. Update the card in the billing portal: `nodus billing portal`.
2. Run `nodus billing auto-recharge --threshold 10 --amount 50` again, or set `enabled: true`.
If your balance runs out while auto-recharge is off, work stops gracefully as described in [When money runs out](../when-money-runs-out/).
## Turn it off
[Section titled “Turn it off”](#turn-it-off)
```bash
nodus billing auto-recharge --off
```
The saved card stays on file for later; remove it in the billing portal.
# Budgets and spending caps
> Put an enforced limit on what an org, a project or a labelled group of work can spend, and cap any single run.
Nodus enforces two kinds of spending limits. Both are checked before any paid work starts and again every few minutes while it runs, so a limit always stops spend instead of reporting it afterwards.
* A **Budget** limits what matching work across your org can spend in a month or in total.
* **`maxCostUSD`** caps one object: a Job, Pipeline, Sweep, TrainingJob, Sandbox, Workspace, Function, Agent or InferenceEndpoint. A cap on a Pipeline or Sweep is shared by every run it creates.
## Create a Budget
[Section titled “Create a Budget”](#create-a-budget)
```yaml
apiVersion: nodus.dev/v1
kind: Budget
metadata:
name: research-monthly
spec:
limitUSD: "500.00"
period: Monthly # a UTC calendar month; Total counts from creation
scope:
project: research # leave out for the whole org
selector:
matchLabels: {team: nlp}
action: Block # Notify only sends notices
thresholds: [50, 80, 100]
notify:
emails: [research-leads@example.com]
```
```bash
nodus apply -f budget.yaml
nodus get budget research-monthly
```
`nodus get budget` shows what is spent, what is held for running work, what remains and the thresholds already notified this period. An org can have up to 50 Budgets.
## What a Budget covers
[Section titled “What a Budget covers”](#what-a-budget-covers)
A Budget counts every hold and charge of work in its project whose labels matched the selector **when the work was created**. Changing a running Job’s labels does not move it out of a Budget. A Budget you create while work is already running starts covering that work within 5 minutes.
## What happens at the limit
[Section titled “What happens at the limit”](#what-happens-at-the-limit)
With `action: Block`:
* New work in scope is refused with `402 BudgetExceeded`, naming the Budget, what is left and what the work needs:
```text
budget "research-monthly" has $0.40 left of $500.00 this month; job "train-a" needs $3.20.
fix: raise spec.limitUSD on budget/research-monthly, or wait until 2026-11-01
```
* Running work in scope stops gracefully, as described in [When money runs out](../when-money-runs-out/), with `Funded=False` and reason `BudgetExceeded`.
* Raising `spec.limitUSD`, switching to `Notify`, deleting the Budget or the start of the next month resumes it.
With `action: Notify`, nothing is refused; you only get the notices.
## Notices
[Section titled “Notices”](#notices)
Each threshold is sent once per period as a `BudgetThreshold` Event, a `budget.threshold` webhook and an email to the billing email, org owners and admins, and the addresses in `spec.notify.emails`. Reaching a `Block` limit also sends `BudgetExceeded` (webhook `budget.exceeded`).
## Cap a single run
[Section titled “Cap a single run”](#cap-a-single-run)
```yaml
kind: Job
spec:
maxCostUSD: "25.00"
```
A Job never holds more than its cap. When its spend reaches the cap, it checkpoints and becomes `Suspended` with reason `MaxCostReached`; raising `maxCostUSD` resumes it from the checkpoint. You can raise a cap at any time, but not lower it.
A Pipeline, Sweep or TrainingJob shares its cap with the Jobs it creates. Each child must fit both its own cap and its parent’s remaining cap, along with every matching Budget. A dry-run reports an insufficient spending limit in `status.estimate.blockingReasons` without reserving funds.
# Buy credits
> Add prepaid credits with a card through Stripe Checkout from the console, the CLI or the API, and add the $20 Indra plan.
Nodus is prepaid. You buy credits, and paid work draws on them through a funded hold placed before it starts. You pay on a Stripe-hosted page, so Nodus never sees your card number. Every purchase is a `TopUp` object that you can list and follow until the credit is on your balance.
## Buy from the console
[Section titled “Buy from the console”](#buy-from-the-console)
Open **Usage & billing**, choose **Add credits**, and pick $10, $20, $50, $100 or a custom amount. The console sends you to Stripe Checkout and brings you back to the billing page, which shows the top-up until it is paid.
## Buy from the CLI
[Section titled “Buy from the CLI”](#buy-from-the-cli)
```bash
nodus billing top-up 20
```
The CLI opens the payment page in your browser, waits until the payment is credited and prints your new balance.
| Flag | Does |
| ------------- | ------------------------------------------------------------------------ |
| `--save-card` | Saves the card for [auto-recharge](../auto-recharge/) |
| `--no-open` | Prints the payment link instead of opening a browser, for a remote shell |
| `--no-wait` | Returns as soon as the payment link exists |
## Buy through the API
[Section titled “Buy through the API”](#buy-through-the-api)
A top-up is an ordinary resource. Create it, then open its `status.checkoutURL`, which appears within about a second:
```yaml
apiVersion: nodus.dev/v1
kind: TopUp
metadata:
name: october-credits
spec:
amountUSD: "50.00"
savePaymentMethod: false # true also saves the card for auto-recharge
```
```bash
nodus apply -f topup.yaml
nodus get topup october-credits -o yaml
```
The REST path is `/apis/nodus.dev/v1/topups`. By default Checkout returns to the console billing page; set `spec.successURL` and `spec.cancelURL` to return somewhere else on the console, or to an `http://127.0.0.1` address for a local tool. A top-up cannot be edited or deleted; an unpaid one expires by itself.
## Limits
[Section titled “Limits”](#limits)
| Limit | Value |
| ----------------------- | -------------------------------------- |
| Amount of one top-up | $5.00 to $1,000.00, in whole cents |
| Top-ups per org per day | $5,000 in total, per UTC day |
| Payment page | Payable for 1 hour after it is created |
An amount outside the range is refused with `422 Invalid` before anything is charged.
## Follow a top-up
[Section titled “Follow a top-up”](#follow-a-top-up)
```bash
nodus get topups -o wide
```
| Phase | Reason | Meaning |
| ----------- | ---------------- | ---------------------------------------------------------------------------------------------------- |
| `Queued` | | Nodus is creating the payment page |
| `Running` | | The payment page is open and waiting for you |
| `Running` | `PaymentPending` | You paid with a method that confirms later, such as a bank debit |
| `Succeeded` | | Paid and credited; `status.creditedUSD` is on your balance and `status.receiptURL` links the receipt |
| `Failed` | `Expired` | Nobody paid within the hour; nothing was charged |
| `Failed` | `PaymentFailed` | The payment did not complete; nothing was credited |
A declined card does not fail the top-up: the payment page stays open so you can try another card. When a payment fails after you submit it, such as a bank debit that does not clear, Nodus emails your billing email.
## When the credit arrives
[Section titled “When the credit arrives”](#when-the-credit-arrives)
Nodus credits each payment once, as soon as Stripe confirms it, which usually takes a few seconds. If a confirmation from Stripe is delayed or lost, Nodus asks Stripe about the payment itself within the hour and again in a nightly check, so a paid top-up is always credited, and never twice.
If your org owes arrears, the top-up pays them first and the rest goes to your balance. Work that stopped because money ran out resumes by itself; see [When money runs out](../when-money-runs-out/).
Stripe emails a receipt for every paid top-up to your billing email. See [Receipts and invoices](../receipts/).
## The Indra plan
[Section titled “The Indra plan”](#the-indra-plan)
The Indra plan costs $20 a month. Each paid month adds a $20 credit that only Indra (`nodus/indra`) requests can spend, at list price. Indra requests use that credit before your other credits, and everything else draws on your regular balance. Unused plan credit expires when the month it belongs to ends; it does not carry over.
Start the plan by opening an Indra plan checkout session (type `ComposerSubscription`) and paying on the page it returns:
```bash
curl -X POST "https://api.nodus-compute.ai/apis/nodus.dev/v1/billingaccounts/default/sessions" \
-H "Authorization: Bearer $NODUS_TOKEN" -H "Idempotency-Key: $(uuidgen)" \
-d '{"type": "ComposerSubscription"}'
```
The response is `{url, expirationTime}`. An org can have one Indra plan; a second request while the plan is active is refused with `409 Conflict`.
* **Renewal.** Stripe charges the card each month and emails the invoice. A new month’s credit is added only once its invoice is paid; while a renewal payment is failing, Stripe retries it and Indra requests draw on your regular credits.
* **Cancel.** Open the billing portal with `nodus billing portal` and cancel the plan there. It ends at the end of the paid month: that month’s credit stays usable until then, and nothing further is charged or added.
# Credits
> The starter credit, credit grants, the order credits are spent in, and when they expire.
Credit grants are credit Nodus gives you, as opposed to credit you buy. Each grant is a `CreditGrant` you can read but not change:
```console
$ nodus get creditgrants
NAME SOURCE AMOUNT REMAINING EXPIRES
grt_01j9x... Starter $30.00 $12.40 2026-10-31
grt_01j9y... Promo $25.00 $25.00 2026-11-30
```
The console lists them under **Usage & billing → Credits**, with what each has spent and when it expires.
## Starter credit
[Section titled “Starter credit”](#starter-credit)
The first org you create gets **$30.00** of starter credit once your email address is verified. It expires 30 days after the org is created. Each person gets starter credit once, however many orgs they create.
Orgs that only have starter credit keep tighter limits until their first top-up: 20 GiB of egress a day, $1 a day of console assistant use and no multi-node runs.
## Where grants come from
[Section titled “Where grants come from”](#where-grants-come-from)
| Source | How you get it |
| ---------- | ------------------------------------------------------------------------ |
| `Starter` | Your first org, once your email is verified |
| `Promo` | [Redeeming a promo code](/docs/guides/billing/promo-codes/) |
| `Admin` | Credit added by Nodus support, for example a pilot or a migrated balance |
| `Goodwill` | Credit from Nodus support for a problem on our side |
Some plan grants pay only for Indra (`nodus/indra`) model calls; they show `nodus/auto`, Indra’s other name, as their scope and are never spent on anything else.
## Spending order
[Section titled “Spending order”](#spending-order)
Charges draw from grants first, soonest expiry first, then from grants that never expire, and only then from purchased credit. You never have to pick which credit pays.
## Expiry
[Section titled “Expiry”](#expiry)
You get an email 7 days and 1 day before a grant expires. At its expiry the unspent remainder leaves your balance at once, even if a run is still using it; the run keeps going on its other credit. Grants are never refunded or turned into cash.
A grant also settles unpaid charges first, the same way a top-up does.
# Partner offers
> Apply an eligible partner offer and understand promotional matching.
Open **Billing → Partner offers**, enter your code and optionally add your company or batch. An organization admin submits the request; a platform administrator verifies eligibility before credit is granted. Organization members can view progress.
For the YC offer, promotional credit reaches **$100 total**, including earlier grants even if spent or expired. A previous $30 starter grant therefore leaves **$70 extra**. You keep the existing **10 GB storage inclusion**.
After approval, earn **10% back on the first $10,000 of eligible usage**, for a maximum **$1,000 bonus**. Only captured usage paid from purchased credit qualifies. Usage paid by starter credit, promotional grants or earned bonuses does not. Earned credit does not expire. Usage corrections adjust the matching benefit, including after an offer is revoked.
You can also use the CLI:
```sh
nodus billing deal
nodus billing deal YC --note 'Company and batch'
```
A request snapshots the terms you applied for. Later offer edits do not change them. Approval keeps the usual card prerequisite and usage prices. If an offer is revoked, future matching stops and existing earned credit remains, subject to corrections.
# Promo codes
> Redeem a promo code for credit, and what each redemption error means.
A promo code adds a [credit grant](/docs/guides/billing/credits/) to your org. Redeem it from the CLI or from **Usage & billing → Credits → Redeem code** in the console:
```console
$ nodus billing redeem LAUNCH25
Redeemed ****CH25: $25.00 of credit, expires 2026-11-30.
```
Codes are not case-sensitive, and spaces around them are ignored. Redeeming needs the `billing:write` permission. The response shows the code masked to its last four characters; Nodus stores only a keyed digest of each code.
## Errors
[Section titled “Errors”](#errors)
| Error | Meaning |
| --------------------- | ---------------------------------------------------------------------------- |
| `404 NotFound` | The code does not exist, or it has been withdrawn |
| `410 Expired` | The code’s last redemption date has passed |
| `409 Conflict` | Your org has already redeemed this code, or the code has been fully redeemed |
| `429 TooManyRequests` | More than 5 attempts in a minute; wait a minute and try again |
Retrying a redemption with the same `Idempotency-Key` returns the first result and never adds credit twice. The CLI and console set the key for you.
## Expiry
[Section titled “Expiry”](#expiry)
A promo grant expires on the date the code sets, counted from when you redeem it. Codes without one never expire.
# Receipts and invoices
> Find the receipt and invoice for every top-up and plan payment, change where receipts are sent, and manage your card and billing details.
Every payment you make to Nodus, whether a top-up you buy, an auto-recharge or an Indra plan month, has a Stripe-hosted invoice with a PDF, and Stripe emails a receipt for it. Nodus does not send postpaid bills: credits are prepaid, and what they paid for is in your [usage records](../usage-and-costs/).
## Where receipts go
[Section titled “Where receipts go”](#where-receipts-go)
Receipts go to the billing email of your org’s `BillingAccount`, which defaults to the email of the owner who created the org. To send them somewhere else, such as a finance mailbox, change `spec.billingEmail`:
```yaml
apiVersion: nodus.dev/v1
kind: BillingAccount
metadata:
name: default
spec:
billingEmail: finance@example.com
```
```bash
nodus apply -f billing.yaml
```
The new address applies to the next payment and is also where Nodus sends money notices, such as a failed auto-recharge, together with the org’s owners and admins.
## Find a receipt
[Section titled “Find a receipt”](#find-a-receipt)
```bash
nodus billing receipts
```
This lists your top-ups, Checkout and auto-recharge alike, with a link to each invoice page and its PDF. The same links are on each `TopUp`:
```bash
nodus get topup october-credits -o yaml
```
| Field | Links to |
| ---------------------- | ------------------------------------------------------------------------- |
| `status.receiptURL` | The Stripe-hosted invoice page, which shows the payment and the card used |
| `status.invoicePDFURL` | The invoice as a PDF |
The links appear a few seconds after the payment succeeds, once Stripe has finalized the invoice.
## The billing portal
[Section titled “The billing portal”](#the-billing-portal)
```bash
nodus billing portal
```
This opens the Stripe billing portal for your org, where you can:
* see and download every invoice and receipt, including Indra plan invoices;
* add, replace or remove the card used for auto-recharge and the Indra plan;
* change the name, address and other details printed on your invoices;
* cancel the Indra plan.
The portal link is valid for 5 minutes; run the command again for a new one. From the API, create a session with `POST /apis/nodus.dev/v1/billingaccounts/default/sessions` and `{"type": "Portal"}`.
Nodus does not charge sales tax or VAT on top-ups.
## Refunds
[Section titled “Refunds”](#refunds)
Refunds cover unspent purchased credit and are made by Nodus support. A refunded top-up shows the amount in `status.refundedUSD`, the credit leaves your balance when the refund is approved, and the refund appears on the top-up’s invoice page. Credit from grants and promo codes is not refundable.
## Questions about a charge
[Section titled “Questions about a charge”](#questions-about-a-charge)
Contact Nodus support before disputing a charge with your bank. While a dispute is open, new work that needs credit is refused with `402 PaymentDisputed`; running work is not stopped. When the dispute closes in your favor the credit returns to your balance.
# Refunds
> How to get unspent purchased credit back to your card, and what can be refunded.
You can get **unspent purchased credit** back to the card that paid for it. Credit grants, including starter credit and promo codes, are never refunded.
## Ask for a refund
[Section titled “Ask for a refund”](#ask-for-a-refund)
Contact support with your org and the top-up you want refunded. Support checks the request and starts the refund; refunds above $500 are approved by a second person at Nodus. You do not need to stop your work first.
The amount leaves your available balance as soon as the refund is approved, so it cannot be spent while the card refund is processing. The top-up shows the refund under **Usage & billing → Receipts & invoices**, and `nodus get transactions` shows a `RefundRequest` and then a `Refund`. Your card is usually credited within 5 to 10 business days. If the card refund fails, a `RefundFailed` transaction returns the amount to your balance.
## What can be refunded
[Section titled “What can be refunded”](#what-can-be-refunded)
* At most what you paid on that top-up, less anything already refunded on it.
* At most your unspent purchased credit at approval: credit already used for runs is not refunded.
* One refund per top-up at a time, in whole cents.
## Disputes
[Section titled “Disputes”](#disputes)
If you dispute a payment with your bank instead, the disputed amount is removed from your balance and new work is paused with `402 PaymentDisputed` until the dispute closes. Running work continues. Contacting support first is usually faster.
# Usage and costs
> See what each run cost, split by project, label, meter and segment, and export usage as CSV.
Every charge Nodus makes is backed by usage records: one line per meter, per object, per window. This page shows how to read them, group them and export them.
## See what a run cost
[Section titled “See what a run cost”](#see-what-a-run-cost)
Each compute object shows its cost so far in `status.cost`:
```bash
nodus get job train-llama -o jsonpath='{.status.cost}'
```
```json
{"totalUSD": "4.182500", "heldUSD": "0.750000", "bySegment": {"bootUSD": "0.121000", "runningUSD": "3.980000", "teardownUSD": "0.081500"}}
```
`heldUSD` is reserved but not yet charged. The total becomes final once the capacity behind the run is confirmed deleted, which can be a minute or two after the run itself ends.
When a run fails because of Nodus, the teardown after the failure and any of its start not yet billed are not charged: `bySegment.coveredByNodus` names those segments, and `nodus describe` shows them on the cost line, for example `$0.00 (covered by Nodus: boot, teardown)`.
## What each segment covers
[Section titled “What each segment covers”](#what-each-segment-covers)
Compute usage from the moment the capacity starts billing until it is confirmed deleted is split into segments:
| Segment | Covers |
| ---------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `Boot` | From the start of billing to the moment your command starts: boot, registration, readiness checks and pulling your image |
| `Running` | Your command running, including reading inputs, lazy image fetches, checkpoint writes and the snapshot when the run stops |
| `Restore` | For a run that resumes from a checkpoint or snapshot: from the start of billing to the moment your command starts again, including the restore. It replaces `Boot` |
| `Teardown` | From the stop to confirmed deletion, plus the capacity’s billing increment (lines with `detail: IncrementRounding`) |
Time between runs that reuse the same capacity is billed on the `warm_idle_seconds` meter and has no segment. When Nodus itself delays a deletion, that time is not charged: it shows as a zero-amount line with `detail: NodusBorne`.
## List and group usage
[Section titled “List and group usage”](#list-and-group-usage)
```bash
nodus get usage --since 2d
nodus get usage --group-by project,meter
nodus get usage --group-by member --since 30d
nodus get usage --group-by label:team,day
nodus get usage --group-by rank,segment --field-selector object.uid=
```
`--group-by` accepts `project`, `kind`, `member`, `meter`, `day`, `segment`, `rank` and `label:`. `member` is the person who started the work, whichever of their API keys they used. Labels are the ones the object carried when it was admitted, so label your work before you submit it:
```yaml
metadata:
name: train-llama
labels:
team: research
```
Grouped totals always add up to the ungrouped total: every line lands in exactly one group, and lines without the grouped value share an empty group.
## Export as CSV
[Section titled “Export as CSV”](#export-as-csv)
```bash
nodus get usage --since 30d -o csv > usage.csv
curl -H "Authorization: Bearer $NODUS_TOKEN" -H "Accept: text/csv" \
"https://api.nodus-compute.ai/v1/usagerecords?since=30d"
```
Amounts are in US dollars with six decimals. `quantity` is in the meter’s unit (seconds, tokens, GiB or characters), and `rate_micros` is the rate in micro-dollars per `rate_basis` units.
## Inference and Indra
[Section titled “Inference and Indra”](#inference-and-indra)
Inference requests are rolled up into one line per model, meter and hour. A request to `nodus/indra` (Indra) shows the model that served it and, beside it, the routing call’s input and output tokens on lines whose SKU ends in `:routing:input` and `:routing:output`. The request is charged once, rounded up to the micro-dollar over all of its lines together.
## Agents
[Section titled “Agents”](#agents)
An Agent’s idle workers bill as warm idle on the Agent. When a run claims a worker, the worker’s time from that moment bills on the AgentRun until the run releases it, so each second of worker time appears once. Model calls a run makes with Nodus’s Claude are billed per request on the AgentRun, with the routing call beside the model that answered; calls made with your own Anthropic key carry no Nodus model charge.
## Storage and egress
[Section titled “Storage and egress”](#storage-and-egress)
Retained storage (checkpoints, Volumes, images you build, and Job and Agent outputs) is sampled every hour per object. The first 10 GB across your organization are included and shared across your objects in proportion to their size; each object’s line shows only its billable part. The hours of a UTC day are charged together shortly after midnight UTC, as one `Storage` transaction.
Egress from your containers and interactive sessions is counted per object per UTC day. The first 10 GiB a day across your organization are included, and the rest is charged the next morning as one `Egress` transaction, split across objects by bytes.
If your balance cannot cover a daily storage or egress charge, the remainder becomes arrears. While arrears are outstanding, new runs and uploads are refused with `ArrearsOutstanding`; your next top-up pays them first.
Storage left in arrears gets an email at 7, 21 and 28 days. At 30 days Nodus proposes deleting the stored data beyond the included 10 GB: finished Jobs’ checkpoints and outputs first, then older Volume revisions, then the latest revisions, oldest first. Nothing is deleted until two Nodus administrators approve the list, and paying your arrears before then cancels it. Deleting data does not clear the arrears.
## Meters reference
[Section titled “Meters reference”](#meters-reference)
`quantity` is in the meter’s unit. `rate` applies to `rateBasis` units of quantity, so a line’s amount is `quantity × rate ÷ rateBasis`, rounded down, except that the last line of a run rounds up to the capacity’s billing increment.
| Meter | Billed for | Unit | Rate is per | `rateBasis` | SKU |
| ------------------------------ | ---------------------------------------------------------------------------------------------------------- | ------------------ | ------------------ | ----------- | -------------------------------------------------------------- |
| `compute_seconds` | Dedicated capacity for Jobs, GPU Sandboxes, GPU Workspaces, Functions and training members, with a segment | seconds | hour | 3 600 | `rented:` |
| `warm_idle_seconds` | Warm capacity kept between runs, idle Function and agent workers | seconds | hour | 3 600 | `rented:` |
| `node_vcpu_seconds` | vCPU on shared nodes (Sandboxes, agent runs), with a segment | milli-vCPU seconds | vCPU-hour | 3 600 000 | `node:vcpu` |
| `node_gib_seconds` | Memory on shared nodes, with a segment | MiB seconds | GiB-hour | 3 686 400 | `node:memory` |
| `node_disk_gib_seconds` | Requested disk above 10 GiB per vCPU on shared nodes | GiB seconds | GiB-hour | 3 600 | `node:disk` |
| `build_vcpu_seconds` | Image builds, with a segment | milli-vCPU seconds | vCPU-hour | 3 600 000 | `build:vcpu` |
| `storage_gb_hours` | Retained storage above 10 GB per organization | MB hours | GB-month (30 days) | 720 000 | `storage:retained` |
| `egress_gib` | Egress above 10 GiB per organization per UTC day; relayed training traffic from the first byte | MiB | GiB | 1 024 | `egress:gib`, `egress:mesh-relay`, `egress:supplier` |
| `inference_input_tokens` | Input tokens | tokens | million tokens | 1 000 000 | `model::input` |
| `inference_output_tokens` | Output tokens, reasoning included | tokens | million tokens | 1 000 000 | `model::output` |
| `inference_cache_read_tokens` | Cached input read | tokens | million tokens | 1 000 000 | `model::cache_read` |
| `inference_cache_write_tokens` | Cache writes, 5-minute or 1-hour | tokens | million tokens | 1 000 000 | `model::cache_write_5m`, `model::cache_write_1h` |
| `inference_audio_seconds` | Transcription and translation, 10 s minimum per request | milliseconds | audio-hour | 3 600 000 | `model::audio` |
| `inference_speech_characters` | Speech synthesis input | characters | million characters | 1 000 000 | `model::speech` |
| `platform_device_hours` | Your own devices assigned to Nodus-scheduled work | device seconds | device-hour | 3 600 | `platform:device-hour` |
| `platform_predict_seconds` | Predict on one of your pools | seconds | 30-day month | 2 592 000 | `platform:predict` |
Indra requests add a routing line beside the model’s lines, with SKU `:routing:input` or `:routing:output`. An `egress:supplier` line passes through an egress charge billed for your dedicated capacity at the same margin as the capacity itself.
# What you pay for
> Every kind of time and cost, who pays for it, and the usage segment it appears under.
You pay for every second a provider bills for a machine Nodus acquired for your work, from the moment billing starts until the machine is confirmed deleted. Nodus pays for everything it chose or caused. Your usage shows each machine’s time by segment, so you can see where every second went:
* **Boot**: from billing start to your command starting.
* **Restore**: a replacement machine’s time until your command starts again after a recovery, including the checkpoint restore.
* **Running**: from your command starting to its stop.
* **Teardown**: from stop to confirmed deletion, plus the provider’s billing increment, rounded once per machine.
Run `nodus get usage --group-by segment` or open **Usage & billing → Usage** in the console to see the split. For a multi-node run, `--group-by rank,segment` itemizes each member.
## The table
[Section titled “The table”](#the-table)
| What | Segment | Who pays |
| --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------- | -------- |
| []()Time from billing start to your command starting: boot, registration, readiness checks, image pull and checkpoint restore | Boot, Restore | You |
| []()A boot that fails on the provider’s side, until the machine is confirmed deleted | Boot | You |
| []()Spare machines Nodus starts to finish your work sooner, machines Nodus releases or replaces on its own decision, machines that fail Nodus’s checks, and failures Nodus causes | None: never on your usage | Nodus |
| []()From your command starting to its stop, including lazy image fetches, checkpoint writes, the snapshot on stop and work lost after a preemption | Running | You |
| []()Downloading inputs after your command starts, and the `initCommand` preflight | Running | You |
| []()Imports, exports and sink loads that run on your org’s capacity | Running | You |
| []()Disk above 10 GiB per vCPU on Nodus nodes | Running | You |
| []()A network partition your work tolerates, until it reconnects or until the old machine is fenced and confirmed deleted | Running, Teardown | You |
| []()The provider’s billing increment, rounded once per machine | Teardown | You |
| []()Warm idle time beyond an increment you already paid for, minimum workers, and idle Function and agent workers | Warm idle | You |
| []()The idle share of Nodus-operated capacity, built into its rates and never metered on its own | Included in rates | You |
| []()From stop to confirmed deletion | Teardown | You |
| []()Time between stop and the delete request beyond 60 seconds when Nodus caused the delay | None: never on your usage | Nodus |
| []()Any cost beyond your tightest limit: `maxCostUSD`, a Budget or your balance | None: never on your usage | Nodus |
| []()A model request whose outcome Nodus cannot confirm, and upstream charges beyond the usage observed on a cut-off stream | None: never on your usage | Nodus |
| []()Console assistant and command generation calls, within the daily cap | None | Nodus |
| []()Grading Sandboxes, evaluations, dataset previews, task manifests and model calls made from inside your containers, under your caps | Running | You |
| []()Mirroring a private image so it can run on hosted containers, and the mirror’s storage | None | Nodus |
| []()Nodus’s own test runs | None | Nodus |
| []()Image builds that fail | Build | You |
| []()Storage above 10 GB per org: Volumes, checkpoints, outputs, images and session state | Storage | You |
| []()Logs | None | Nodus |
| []()Egress within 10 GiB per org per day | None | Nodus |
| []()Egress beyond the included 10 GiB a day, up to your quota, and relayed traffic between multi-node members | Egress | You |
| []()Your own pool hosts and cloud accounts | None | Free |
| []()Pool devices running work Nodus schedules, Predict per pool, and capacity Nodus acquires for you | Platform | You |
| []()Differences between a provider’s invoice and what Nodus metered | None | Nodus |
This table is checked against the billing reference model on every change, so the rows here are the rows Nodus charges by.
## Limits stop work before it overruns
[Section titled “Limits stop work before it overruns”](#limits-stop-work-before-it-overruns)
Every paid action holds funds before it starts, and work stops gracefully when a hold cannot be renewed. If a charge would still go past your tightest limit, Nodus absorbs the difference: your balance never pays more than the limit you set. See [How billing works](/docs/concepts/billing/) for holds, captures and the low-balance stop.
# When money runs out
> How Nodus warns you, stops work gracefully and resumes it when credits, a raised Budget or a raised cap return.
Credits are prepaid. Before paid work starts, Nodus places a funded hold on your balance, and it renews that hold every 5 minutes while the work runs. When a renewal cannot be funded, from your balance, a Budget or a `maxCostUSD` cap, the work stops gracefully inside time that is already paid for. Nothing runs past what you funded, and nothing is lost that the work saved.
## Before you start: fail fast
[Section titled “Before you start: fail fast”](#before-you-start-fail-fast)
A create that cannot be funded is refused at once with the amounts and the fix:
```text
402 InsufficientCredits
job "train-a" needs a $3.20 hold to start (released when it ends); available $1.10.
fix: nodus billing top-up 20, or lower spec.maxCostUSD
```
Other refusals are `BudgetExceeded` (a Budget or a cap has too little left), `ArrearsOutstanding` (unpaid charges) and `PaymentDisputed`. A distributed Job checks all of its nodes together: it starts with every node funded or not at all.
## Warnings
[Section titled “Warnings”](#warnings)
When your available balance falls below $5, or below 20 % of what the next renewal of your running work needs, you get a `LowBalance` Event, a `billing.low_balance` webhook, at most one email a day and a console banner. Running work whose funds will run out soon still shows `Funded=True`, and the condition’s message gives the time it will stop. A top-up, a credit grant or an auto-recharge before that time keeps it running.
When a charge is larger than what is left, the rest becomes unpaid charges. You get an `ArrearsPosted` Event, a `billing.arrears_posted` webhook and at most one email a day, and new work waits until a top-up or credit grant pays them.
## How each kind stops
[Section titled “How each kind stops”](#how-each-kind-stops)
| Kind | What happens | State | Resumes |
| ------------------------------ | ----------------------------------------------------------------------------------- | -------------------------------------------------------------------------- | --------------------------------------------------------- |
| Job (checkpointed) | Takes an urgent checkpoint, then stops | `Suspended`, `Funded=False` | Automatically once funded, from the checkpoint |
| Job (restartable or ephemeral) | Stops at the funded edge | `Suspended`, `Funded=False` | Automatically; ephemeral Jobs restart from the start |
| Job at `maxCostUSD` | Checkpoints, then stops | `Suspended`, reason `MaxCostReached` | When you raise `maxCostUSD` |
| Distributed Job | Every node checkpoints together, then all stop | `Suspended`, `Funded=False` | Automatically, with all nodes funded |
| Sandbox | Snapshots its filesystem and state, then stops | `Stopped`, `stopReason: InsufficientCredits`; `spec.state` stays `Running` | On the next exec or connect once funded |
| Workspace | Saves the home volume, then stops | `Stopped` with the money reason | On `nodus start`, SSH connect or the schedule once funded |
| AgentRun | Finishes or interrupts the current step | `Waiting` for funds | Automatically |
| Agent idle workers | Scale to zero | `Funded=False` on the Agent | Automatically |
| Function | Stops taking calls, finishes in-flight calls, scales to zero | New calls queue with `Funded=False` | Automatically |
| TrainingJob | Its Jobs suspend; grading Sandboxes are deleted | `Suspended` | Automatically |
| Image build | The build stops | Image `Pending`, `Funded=False` | Automatically rebuilds |
| Volume import or export | The transfer stops | `ImportFailed(InsufficientCredits)` | `nodus request reimport volume/` |
| Inference | New requests return an OpenAI- or Anthropic-shaped `402`; in-flight requests finish | — | Immediately |
Nodus never changes the `spec.state` you set. It records the stop in the object’s phase and `Funded` condition, emits `FundingLost`, and emits `FundingRestored` when it resumes. You get at most one email an hour listing stopped work.
## Getting going again
[Section titled “Getting going again”](#getting-going-again)
```bash
nodus billing top-up 20
nodus get jobs --watch
```
A top-up resumes waiting work oldest first, as each hold fits. Work stopped by a Budget resumes when you raise the Budget’s limit, switch it to `Notify`, delete it or the next month starts; work stopped at `maxCostUSD` resumes when you raise the cap.
# Checkpoints
> Keep a Job's progress through suspends, stops and lost capacity by saving it to the state directory, and answer checkpoint requests so every save is consistent.
A checkpoint is a copy of your Job’s **state directory** that Nodus takes while the Job runs. When the Job is suspended, or the capacity under it is reclaimed, the next attempt starts with that directory restored and your program carries on from what it saved. Nodus decides when to checkpoint and where to store it; your program decides what to write and how to load it.
Checkpoints hold files, not memory. A restored attempt starts your command again from the beginning, with the state directory as it was at the last checkpoint, so the program must read its own progress back.
## Save your progress to the state directory
[Section titled “Save your progress to the state directory”](#save-your-progress-to-the-state-directory)
Write everything you need to continue, such as model weights, optimizer state and the current step, under `/nodus/state`. The path is also in `NODUS_STATE_DIR` (`NODUS_CHECKPOINT_DIR` is the older name for the same directory). On start, load what is there:
examples/checkpoints/resume/train.py
```python
"""A training loop that saves its progress where Nodus checkpoints it and resumes from it.
It uses only the standard library, so it runs in any image. With the Nodus SDK installed,
`nodus.state_dir()`, `nodus.checkpoint.on_request()` and `nodus.restored()` do the same work.
"""
import json
import os
import socket
import threading
import time
from pathlib import Path
STATE_DIR = Path(os.environ.get("NODUS_STATE_DIR", "/nodus/state"))
STATE = STATE_DIR / "progress.json"
STEPS = int(os.environ.get("STEPS", "120"))
SAVE_EVERY = 20 # like Hugging Face Trainer's save_steps: a regular save even when Nodus does not ask
pending = threading.Event() # set while Nodus waits for a consistent checkpoint
request_seq = None
events = None
def save(step):
"""Write the state atomically, so a snapshot never sees a half-written file."""
STATE_DIR.mkdir(parents=True, exist_ok=True)
tmp = STATE.with_suffix(".tmp")
tmp.write_text(json.dumps({"step": step}))
os.replace(tmp, STATE)
def listen():
"""Subscribe to checkpoint requests on the events socket and flag each one for the loop."""
global events, request_seq
path = os.environ.get("NODUS_EVENTS_SOCKET", "/run/nodus/events.sock")
try:
events = socket.socket(socket.AF_UNIX, socket.SOCK_STREAM)
events.connect(path)
except OSError:
return # outside Nodus there is nobody to ask for checkpoints
events.sendall(b'{"type":"checkpoint.subscribe"}\n')
for line in events.makefile("r"):
message = json.loads(line)
if message.get("type") == "checkpoint.request":
request_seq = message["seq"]
pending.set()
def ack():
"""Tell Nodus the files in the state directory are complete."""
events.sendall((json.dumps({"type": "checkpoint.ack", "seq": request_seq}) + "\n").encode())
pending.clear()
def main():
start = 0
if STATE.exists():
start = json.loads(STATE.read_text())["step"]
print(f"resumed from step {start}", flush=True)
threading.Thread(target=listen, daemon=True).start()
for step in range(start + 1, STEPS + 1):
time.sleep(1) # one step of work
print(f"step {step}/{STEPS}", flush=True)
if pending.is_set():
save(step)
ack()
elif step % SAVE_EVERY == 0:
save(step)
print("training complete", flush=True)
if __name__ == "__main__":
main()
```
Two habits keep checkpoints useful:
* **Write atomically.** Write to a temporary file and rename it over the old one, as `save()` does. A checkpoint can then never capture a half-written file.
* **Keep outputs separate.** Results you want to download go to `/nodus/outputs`. The state directory is recovery state: it is restored into the next attempt, not offered as a download.
`NODUS_RESTORED=1` is set in an attempt that started from a checkpoint, if your program wants to log the difference. An empty state directory never replaces an earlier checkpoint that had files in it, so an attempt that fails before it saves anything cannot erase progress.
## Choose what is saved
[Section titled “Choose what is saved”](#choose-what-is-saved)
`recovery.checkpoint` in the Job spec controls the checkpoint. The defaults suit most programs:
```yaml
recovery:
continuity: Checkpointed # restore the latest checkpoint into each new attempt
checkpoint:
paths: [/nodus/state] # up to 64 absolute paths, all saved in one checkpoint
interval: auto # or a fixed interval from 1m to 6h
maxSize: 1Ti # larger checkpoints fail with CheckpointTooLarge
retainAfterFinish: 168h # after the Job finishes, keep only the final checkpoint
```
On the command line, `nodus run --checkpoint /nodus/state` sets `paths`. Saving a whole folder such as your working directory is possible by listing it, but it makes every checkpoint larger and slower; list only what you need to continue.
`continuity: Restartable` is the lighter choice for programs that track their position as a counter: Nodus keeps `NODUS_CURSOR_COMPLETED` and `NODUS_CURSOR_TOTAL` from your progress reports and skips the file copy. `continuity: Ephemeral` starts every attempt from scratch.
## When Nodus checkpoints
[Section titled “When Nodus checkpoints”](#when-nodus-checkpoints)
With `interval: auto`, Nodus sets the cadence from how often the capacity your Job runs on is interrupted and how long a checkpoint takes to save. A Job gets at least four checkpoints over its expected run time, and checkpointing takes no more than about a tenth of it. Capacity that is rarely interrupted is checkpointed less often.
Nodus also checkpoints, whatever the interval:
* when you run `nodus suspend job/NAME`, before the compute is released, so `nodus resume` continues from there;
* when the capacity gives notice that it is about to be reclaimed, so the replacement attempt loses as little work as possible.
`nodus describe job/NAME` shows the latest checkpoint, and `GET …/jobs/NAME/checkpoints` lists each one with its sequence number, attempt, time, size and file count.
## Answer checkpoint requests
[Section titled “Answer checkpoint requests”](#answer-checkpoint-requests)
A checkpoint taken while your program is halfway through writing its files would restore a broken state. To avoid that, a program can ask to be told before each checkpoint and say when its files are complete. This is the **request/ack handshake**, and the example above implements it with the standard library.
It runs over the events socket at `/run/nodus/events.sock` (`NODUS_EVENTS_SOCKET`), one JSON object per line:
1. Your program connects and sends `{"type": "checkpoint.subscribe"}` once.
2. Before each checkpoint, Nodus sends `{"type": "checkpoint.request", "seq": 7, "urgent": false}`.
3. Your program finishes the current step, writes its state and replies `{"type": "checkpoint.ack", "seq": 7}`.
4. Nodus copies the state directory, then your program carries on. It does not need to pause while the copy runs.
`urgent: true` means the capacity is about to go away: save at the next safe point and skip optional work. Nodus waits for the ack for as long as the shutdown allows and then takes the checkpoint anyway, so a program that hangs cannot block a suspend.
You choose whether to use the handshake with `recovery.checkpoint.integration`:
| Value | Behaviour |
| ---------------- | --------------------------------------------------------------------------------------------------------------- |
| `Auto` (default) | Use the handshake when the program subscribes; otherwise checkpoint without asking |
| `None` | Never ask; checkpoint the paths as they are |
| `HFTrainer` | Answer requests from inside Hugging Face Trainer, with no change to your image ([below](#hugging-face-trainer)) |
With the Python SDK installed (`pip install nodus-compute`), `nodus.checkpoint.on_request(save)` registers a callback and sends the ack after it returns, `nodus.checkpoint.requested()` lets a loop poll instead, and `nodus.state_dir()` returns the directory. All of them do nothing outside Nodus, so the same script runs on your laptop.
## Hugging Face Trainer
[Section titled “Hugging Face Trainer”](#hugging-face-trainer)
Trainer already saves and resumes; point it at the state directory and it works with Nodus checkpoints:
* set `output_dir` to the state directory (`os.environ["NODUS_STATE_DIR"]`), or a folder inside it;
* set `save_steps` to how often Trainer saves on its own, and `save_total_limit` (for example `2`) so old checkpoints do not fill the directory;
* call `trainer.train(resume_from_checkpoint=True)` when `NODUS_RESTORED` is `1`, and `trainer.train()` on a first start, because Trainer refuses to resume from an empty directory.
To also save when Nodus asks, add `nodus.checkpoint.HFTrainerCallback()` to the Trainer’s callbacks, or set `integration: HFTrainer` and Nodus registers the same callback for you, even in an image without the SDK. Trainer then saves at the end of the current step and Nodus checkpoints once the save is written.
## Gang checkpoints (Beta)
[Section titled “Gang checkpoints (Beta)”](#gang-checkpoints-beta)
Beta
Multi-node Jobs (`distributed`) are Beta. Their checkpoints use a different format.
A Job with `distributed` set defaults to `recovery.checkpoint.format: Dcp`. Every rank writes its shard with `torch.distributed.checkpoint` to `$NODUS_CHECKPOINT_URI`, and rank 0 commits it, which `nodus.checkpoint.dcp.save` and `.load` do for you. After a restart, `$NODUS_RESTORE_URI` names the latest committed checkpoint. A rank that lost its place in the gang cannot commit, so a restored gang always loads a checkpoint every rank finished.
## Run the example
[Section titled “Run the example”](#run-the-example)
The example suspends the Job mid-run and resumes it; the second attempt prints `resumed from step N` and finishes the remaining steps:
```console
$ cd examples/checkpoints/resume
$ nodus run --name checkpoints-resume --cpu 2 --checkpoint /nodus/state -d -- python train.py
$ nodus suspend job/checkpoints-resume
$ nodus resume job/checkpoints-resume
$ nodus logs -f job/checkpoints-resume
resumed from step 12
step 13/120
…
```
Checkpoints are deleted with their Job. After a Job finishes, Nodus keeps only its final checkpoint once `retainAfterFinish` (seven days by default) has passed.
# Which Nodus compute feature should I use?
> Compare Jobs, Workspaces, Sandboxes and Functions. Choose the right way to run code, call models, run agents or train a model on Nodus.
Use a **Job** when you have a command that should run to completion. Use a **Workspace** for interactive GPU or CPU development, a **Sandbox** for isolated CPU code execution, and a **Function** to call Python remotely or in parallel. Nodus also provides hosted inference, managed agents and training workflows.
## Choose by the work you need to do
[Section titled “Choose by the work you need to do”](#choose-by-the-work-you-need-to-do)
| I need to… | Start with | Why it fits |
| ----------------------------------------------------- | ------------------------------------------- | --------------------------------------------------------------------------------- |
| Run a script, batch task or existing training command | [Jobs](/docs/guides/jobs/) | Runs a command to completion with logs, outputs and resource and cost limits |
| Develop with SSH, VS Code or JupyterLab | [Workspaces](/docs/guides/workspaces/) | An interactive GPU or CPU machine with a saved home directory |
| Execute agent-generated or untrusted code | [Sandboxes](/docs/guides/sandboxes/) | An isolated CPU container driven by command and file requests, with idle stopping |
| Call Python remotely or map it over many inputs | [Functions](/docs/guides/functions/) | Remote calls with workers that scale with demand |
| Call a hosted language model | [Inference (Beta)](/docs/guides/inference/) | An OpenAI-compatible model API billed per token |
| Run an agent conversation with shell tools | [Agents (Beta)](/docs/guides/agents/) | An AgentRun with its own sandbox and recorded execution history |
| Fine-tune or train through managed runtimes | [Training (Beta)](/docs/guides/training/) | A TrainingJob defines the model, data, runtime and training parameters |
For a first run, follow [your first job](/docs/getting-started/). It walks through signing in, running a command, following logs and downloading results.
## What is the difference between a Job and a Function?
[Section titled “What is the difference between a Job and a Function?”](#what-is-the-difference-between-a-job-and-a-function)
A Job runs an executable command and finishes when that command exits. A Function is a Python callable deployed in an App; your code invokes it with `.remote()`, `.map()` or `.spawn()`. Choose Jobs for an existing script or batch process. Choose Functions when remote calls should be part of your Python application.
Both can use GPU or CPU workers. Function workers can stay warm between calls; read [Function billing](/docs/guides/functions/billing/) before choosing idle and scaling settings.
## What is the difference between a Workspace and a Sandbox?
[Section titled “What is the difference between a Workspace and a Sandbox?”](#what-is-the-difference-between-a-workspace-and-a-sandbox)
A Workspace is for interactive development with SSH, VS Code or JupyterLab on GPU or CPU compute. Its home directory is backed by a Volume. A Sandbox is an isolated CPU container for programmatic commands and file operations, including code produced by an agent. Its network is closed unless you open it, and it can stop after an idle period.
Read [Workspace storage](/docs/guides/workspaces/) and [Sandbox isolation](/docs/concepts/sandboxes-isolation/) before deciding what state and access your task needs.
## Do I need Agents to use Claude Code, Codex or Cursor?
[Section titled “Do I need Agents to use Claude Code, Codex or Cursor?”](#do-i-need-agents-to-use-claude-code-codex-or-cursor)
No. Connect your existing coding agent to Nodus through [MCP](/docs/guides/mcp/) using the [client setup guide](/connect/). It can work with Nodus resources through that connection. The Agents feature is for running an agent conversation inside Nodus itself.
## What survives an interruption?
[Section titled “What survives an interruption?”](#what-survives-an-interruption)
Recovery depends on the feature and its configuration. For checkpointed Jobs, your program writes and loads its own files in `NODUS_CHECKPOINT_DIR`; Nodus saves and restores that directory. This does not restore arbitrary process or GPU memory. Restartable Jobs start over, while ephemeral Jobs keep no state.
Read [Checkpoints](/docs/guides/checkpoints/) for application recovery, [Volumes](/docs/guides/volumes/) for persistent files and [Outputs](/docs/guides/outputs/) for results you need to download.
## How do I control spending?
[Section titled “How do I control spending?”](#how-do-i-control-spending)
Check the estimate before starting work, set the resource’s supported cost and time limits, and configure [project budgets](/docs/guides/billing/budgets/). Compute usage, saved storage and model tokens have different billing rules; the [billing guide](/docs/guides/billing/) explains them. Use the [pricing reference](/docs/reference/pricing/) for published amounts instead of copying prices from an old example.
## Where should an agent look up exact syntax?
[Section titled “Where should an agent look up exact syntax?”](#where-should-an-agent-look-up-exact-syntax)
Use the [CLI reference](/docs/reference/cli/) for commands, [Python SDK reference](/docs/reference/python/) for signatures and [OpenAPI](/docs/openapi.json) for v1 request and response fields. Beta resources use the [v1beta1 contract](/docs/openapi-v1beta1.json).
Every authored guide has a Markdown version with its examples. The [agent guide](/docs/for-agents/) links a lightweight page manifest and focused topic bundles for retrieval.
# Connections
> Connect databases, S3 buckets, Weights & Biases and GitHub once, verified, and use them from Jobs, Sandboxes, imports and outputs.
A Connection is an external system that Nodus talks to for you: a Postgres database for query imports and outputs, an S3 bucket for inputs, a Weights & Biases project for live run tracking, or your GitHub repositories for private sources. Its credentials live in a [Secret](/docs/guides/secrets/); the Connection says what they are for, and Nodus checks that they work before anything uses them.
## Create a Connection
[Section titled “Create a Connection”](#create-a-connection)
Put the credentials in a Secret, then point the Connection at it:
```console
nodus secret create analytics-db --from-literal DATABASE_URL=postgres://reader:...@db.example.com:5432/app
```
connection.yaml
```yaml
apiVersion: nodus.dev/v1
kind: Connection
metadata:
name: ex-connections-postgres
spec:
type: Postgres
secret: ex-connections-postgres # holds DATABASE_URL
scope: Read # verified: the role can read, and is not asked to write
```
```console
nodus apply -f connection.yaml
nodus wait connection/ex-connections-postgres --for=condition=Verified
nodus describe connection/ex-connections-postgres
```
| `type` | Secret keys | Settings | Verified by |
| ------------------ | -------------------------------------------------------------- | ---------------------------------------------------- | -------------------------------------------------------------------- |
| `Postgres` | `DATABASE_URL` | `scope`: `Read`, `Write` or `ReadWrite` | Connecting, a query, and for write scopes the right to create tables |
| `Neon`, `Supabase` | `DATABASE_URL` | `scope`; `neon.branch` | The same checks, on the provider’s own host |
| `S3` | `ROLE_ARN`, or `AWS_ACCESS_KEY_ID` and `AWS_SECRET_ACCESS_KEY` | `s3.bucket`, `s3.region`, `s3.prefix`, `s3.endpoint` | Reaching the bucket with short-lived credentials |
| `WandB` | `WANDB_API_KEY` | `wandb.entity`, `wandb.project`, `wandb.live` | The key’s access to the entity |
Prefer an IAM role (`ROLE_ARN`) for S3: Nodus assumes it for minutes at a time, so no long-lived key is stored. Every S3 Connection has its own external id in `status.externalId`, and Nodus sends exactly that id each time it assumes the role, so a role your trust policy grants to one Connection can’t be used through any other Connection or organization. Create the Connection, read the id, and require it in the role’s trust policy:
```console
nodus get connection/datasets -o jsonpath='{.status.externalId}'
```
```json
{
"Effect": "Allow",
"Principal": { "AWS": "" },
"Action": "sts:AssumeRole",
"Condition": { "StringEquals": { "sts:ExternalId": "" } }
}
```
The Connection shows `Failed` until the trust policy names the id; the next check, within the hour, turns it `Ready`. You don’t need `EXTERNAL_ID` in the Secret. If you set it, it must equal `status.externalId`, and any other value is refused.
A Connection is `Ready` once verified, and `status.egressHosts` lists the hosts it may reach. Nodus checks it again every day, and every hour after a failure. When a check fails the Connection shows `Failed`, the `Verified` condition says why, and anything that needs it fails with `ConnectionNotReady` until you fix the Secret.
## Use it
[Section titled “Use it”](#use-it)
```yaml
spec:
connections: [tracking] # injects WANDB_* and allows its hosts through the egress policy
inputs:
- name: shards
bucket: {uri: s3://acme-data/shards/, connection: datasets}
```
A `WandB` Connection with `wandb.live: true` sets `WANDB_API_KEY`, `WANDB_ENTITY`, `WANDB_PROJECT`, `WANDB_RUN_GROUP`, `WANDB_NAME` and `WANDB_RUN_ID` for the Job, so `wandb.init()` needs no arguments. Each index of the Job logs to one run, and a run that Nodus restarts after lost capacity resumes it rather than starting a new one. A Job with one index shows the run’s page in `status.links`. A variable you set yourself keeps your value; setting `WANDB_RUN_ID`, `WANDB_ENTITY` or `WANDB_PROJECT` yourself means Nodus doesn’t link the run. A Job that names a Connection that is missing or not `Ready` runs without it and gets a `ConnectionNotReady` Event.
## Connect GitHub
[Section titled “Connect GitHub”](#connect-github)
A `GitHub` Connection installs the Nodus GitHub App on your account or organization. It needs no Secret:
```console
nodus create connection github --type GitHub
nodus get conn github -o jsonpath='{.status.installURL}'
```
Open the link, choose the repositories to share and install the App. GitHub returns you to Nodus, which checks that you can see the installation, and the Connection becomes `Ready` with `status.github` naming the account and repositories. `GET …/connections/github/repositories?limit=100` pages through all of them.
A Sandbox’s `init.git` can then clone a private repository: Nodus resolves the branch or tag to a commit when you create the Sandbox and clones that commit. The clone uses a short-lived token that can only read that one repository, and the token never appears in the Sandbox’s environment or files. If someone uninstalls the App, the Connection shows `Failed` with `InstallationRemoved` and a fresh `status.installURL` to install it again.
A Volume can import from a bucket (`source.s3`) or a query (`source.connectionQuery`); see [Volumes](/docs/guides/volumes/#import-from-elsewhere). A Job or Sandbox reaches only the hosts its Connections were verified against.
# Use the console
> Sign in, find your work, follow it live and copy the matching CLI command from any page.
The console at `console.nodus-compute.ai` shows the same objects the CLI and the Python SDK work with. Every page has the command that does the same thing, so you can move between them at any time.
## Sign in
[Section titled “Sign in”](#sign-in)
1. Open the console and sign in with your email and password, or with GitHub or Google.
2. New accounts confirm their email with a 6-digit code.
3. If someone invited you, the console shows the invite first. Accept it to join their organization.
4. Otherwise create your organization. Its URL name (for example `research`) appears in every link.
Your first organization receives starter credit once your email is verified.
## Sign in the CLI from the browser
[Section titled “Sign in the CLI from the browser”](#sign-in-the-cli-from-the-browser)
Run `nodus login`. The CLI opens the console, which asks which organizations the CLI may use and names the key it creates, for example `cli-laptop-2026-09-30`. Choose **Authorize** and return to your terminal. On a machine without a browser, run `nodus login --device` and enter the code it prints on the console’s device page.
If you have no organization yet, authorizing creates one for you. If someone invited you, both pages show the invite first: accept it, and the CLI signs in to their organization.
## Find your work
[Section titled “Find your work”](#find-your-work)
* The URL always names the organization and project: `/research/default/jobs/lora-fine-tune/logs`. Links you share open in the same place for anyone with access, and two browser tabs can work in two organizations.
* Switch organization or project from the top bar.
* Press `Cmd+K` (or `Ctrl+K`) to search objects, pages and actions. Type `job/` or `sb/` and a name to jump straight to an object.
* Press `?` for every keyboard shortcut. `g` then a letter goes to a section, for example `g b` for Billing.
## Follow runs live
[Section titled “Follow runs live”](#follow-runs-live)
Lists and detail pages update as things change, without reloading. The note **Live · 12s ago** shows the time since the last update. Choose it and **Pause live updates** to freeze what you are reading; **Resume** applies the changes at once. Running work shows its cost so far, marked with `≈` between updates.
## Copy the command
[Section titled “Copy the command”](#copy-the-command)
The terminal icon in each page header copies the command for what you are looking at, with the project and organization spelled out:
```sh
nodus describe job/lora-fine-tune -p default --org research
```
**View YAML** in the actions menu shows the live object as a manifest you can download and apply elsewhere.
## Create something
[Section titled “Create something”](#create-something)
Create pages show four tabs: **Form**, **YAML**, **CLI** and **Python**. They always show the same manifest. Before you launch, the estimate shows the expected cost range, the hourly rate, the first hold and how long the price holds. Launch uses exactly that estimate; if prices move first, the console estimates again and asks you to confirm. Your form is saved as a draft in this browser until you launch or discard it.
If the network drops while you launch, the console says it could not confirm the result. **Check status** looks for the object by name, and **Retry** is safe: it can never create a second copy.
## When credits run out
[Section titled “When credits run out”](#when-credits-run-out)
If your balance cannot fund the next renewal, running work stops at its funded edge. The project Overview then shows **Paused for funds**, grouped by what adding credits does:
* **Resumes automatically**: jobs and agent runs continue by themselves once a hold fits.
* **Needs a start**: workspaces stay stopped until you choose **Start** (or **Start all**).
* **Wakes on next use**: sandboxes start again on the next command, file request or preview.
* **Rejecting requests**: inference endpoints answer with `402` until funded.
* **Failed imports**: choose **Retry import** on the volume.
* **Needs a higher limit**: work that reached its own spend limit resumes only when you raise that limit.
Choose **Add credits** on the panel to top up.
## Appearance
[Section titled “Appearance”](#appearance)
Choose **System**, **Light** or **Dark**, and **Comfortable** or **Compact** density, from the account menu. The theme follows you across the console and the docs.
# Environments
> Task sets with graders for reinforcement learning and evaluation, from the Nodus catalog or your own image.
An Environment is a versioned set of tasks with a grader: a prompt for each task, and a program that decides whether a completion is correct. TrainingJobs use Environments for reinforcement learning and evaluation, and agent evaluations use the same ones. Environments are generally available (`nodus.dev/v1`).
## Browse the catalog
[Section titled “Browse the catalog”](#browse-the-catalog)
```console
$ nodus get environments -n nodus
NAME VERSION CATEGORY MODES PHASE
graph-coloring 1.0.0 Reasoning Train, Evaluate Ready
arithmetic 2.0.0 Math Train, Evaluate Ready
gsm8k 1.0.0 Math Train, Evaluate Ready
reasoning-gym 1.0.0 Reasoning Train, Evaluate Ready
python-functions 1.0.0 Code Train, Evaluate Ready
$ nodus describe environment/graph-coloring -n nodus
```
`describe` shows the summary, licenses, the size of each split, a sample task, the graders and the examples: ready TrainingJob templates with measured results. The console’s Training page lists the same catalog, and “Run example” starts a TrainingJob from one; in Python, `nodus.TrainingJob.from_example("nodus/graph-coloring@1.0.0")` does the same.
## Use one in a TrainingJob
[Section titled “Use one in a TrainingJob”](#use-one-in-a-trainingjob)
Name the environment and version, and how many tasks to train and evaluate on:
```yaml
spec:
runtime: nodus/grpo-lora
environment:
name: nodus/graph-coloring@1.0.0
trainTasks: 50 # from the train split
heldOutTasks: 64 # from the test split, never trained on
seed: 42
```
The same name, version and seed always give the same tasks in the same order, so two runs are comparable. The train and test splits never share a task.
## How grading works
[Section titled “How grading works”](#how-grading-works)
* Tasks carry a prompt and non-secret metadata only. Expected answers stay with the grader; the trainer, the model and your code never see them.
* Completions are graded by Nodus, not by the trainer. Program graders run in grading Sandboxes that Nodus creates for your TrainingJob or agent evaluation: no network access, deleted after five minutes idle, at most `grading.maxParallel` at once. Code from a completion runs as a separate user that cannot read the answers. `ExactMatch` graders compare the completion with the answer inside Nodus, without a Sandbox.
* Every verdict is `Correct`, `Incorrect`, `InvalidOutput` or `InfrastructureFailure`, with a reward and evidence recorded by the host that ran the grader: the process, its exit code, its duration and a hash of its output.
* `InvalidOutput` is a completion the grader cannot parse, such as an answer without the expected tag. It is a failed task with a reward of 0.
* `InfrastructureFailure` means the grader could not run. It carries no reward and does not count toward a pass rate, so an outage never looks like a wrong answer.
* Grading Sandboxes are billed to the TrainingJob and count toward its `maxCostUSD`. They are deleted when the run finishes or is suspended.
## Train on a reward function in one call
[Section titled “Train on a reward function in one call”](#train-on-a-reward-function-in-one-call)
When your tasks and reward are in Python, `rl.train` is all you need. To see it work first, run the example that ships with the SDK:
```sh
python -m nodus.examples.rl
```
It is one file, `nodus/examples/rl.py`, and it is the template for your own task:
```python
import re
from nodus.recipes import rl
ANSWER = re.compile(r"\s*([A-Za-z]+)\s*")
tasks = [(f"Spell the word backwards inside .\n\nWord: {w}", w[::-1]) for w in WORDS]
def reward(completion, answer):
found = ANSWER.findall(completion)
return None if not found else float(found[-1].lower() == answer)
run = rl.train(tasks, reward, max_cost=2)
run.watch() # each stage, then the reward, loss and KL of every step
run.outputs.download("./outputs") # the LoRA adapter and the before-and-after comparison
```
`run.tasks(phase="Evaluation", outcome="Failed")` lists the graded tasks the run has scored so far, each with its `taskId`, `reward` and outcome, so you can read which tasks the trained model still gets wrong.
* **Tasks** are `(prompt, answer)` pairs or `{"prompt", "answer", "metadata"}` dicts. Only `reward` sees the answer.
* **Model.** The base model is `Qwen/Qwen3-0.6B` unless you pass `model=`, for example `model="Qwen/Qwen3-4B"`.
* **Held-out tasks.** A fifth of the tasks, and at least 16, is held out. The model is graded on them before and after training. Pass `test_tasks=` to choose them yourself.
* **What gets packaged.** The reward goes with the lines of its own file that it uses: imports, constants and helper functions. A name it cannot take along fails before anything is built.
* **Checked before it runs.** `rl.train` grades one task’s answer and an empty reply on your machine first. A reward that raises, or returns something other than a number, a bool or `None`, fails there instead of on a GPU. Return `None` or `0` when a reply holds no answer.
* **Reuse.** The same code and tasks reuse the same Environment, so a second run starts at once.
* **No network.** The reward runs in Nodus’s grader, which has no network access. Pass `pip=["package==1.2.3"]` for packages it imports.
* **How it learns.** Each step samples `group_size=8` replies to each of `groups_per_step=4` tasks and scores every reply against the others in its group. Replies are capped at `max_tokens=256`, and the learning rate is `learning_rate=4e-5`. You can change any of them.
* **Format rewards.** To reward the shape of an answer as well as its value, fold both into the reward, for example `return float(correct) - 0.1 * (not well_formed)`.
* **Other options.** `steps`, `gpu`, `max_cost`, `lora` and the other `grpo_lora` parameters are optional. The run appears under **Runs › Post-training** in the console, with its curves, results and cost.
To train on a catalog Environment, pass its name instead of tasks and a reward:
```python
run = rl.train("nodus/gsm8k@1.0.0", max_cost=5)
```
## Bring your own environment
[Section titled “Bring your own environment”](#bring-your-own-environment)
Your own tasks and reward are a Python module with two functions. `tasks(split)` lists the prompts of the `train` or `test` split with their answers, and `reward(completion, answer)` scores one completion:
```python
def tasks(split):
return [{"prompt": "Spell the word backwards...\n\nWord: valley", "answer": "yellav"}, ...]
def reward(completion, answer):
found = ANSWER.findall(completion)
if not found:
return None # no answer in the completion: InvalidOutput
return 1.0 if found[-1].lower() == answer else 0.0
```
A reward of 1 (or `True`) is `Correct`, anything lower `Incorrect`, and `None` `InvalidOutput`. If `reward` raises, the verdict is `InfrastructureFailure`, never a reward of 0. Answers reach only `reward`; the trainer and the model see the prompt and the optional `metadata`. A test prompt never appears in train, even if your lists repeat it.
The quickest way in needs no Docker: publish the module as a pip package whose `load_environment()` returns an object (or the module itself) with `tasks` and `reward`, and name the exact version. Nodus builds it into an image when you apply the Environment; if the package does not load, the build fails and the Image `env--` shows the step and its log:
```yaml
apiVersion: nodus.dev/v1
kind: Environment
metadata:
name: reverse-words
spec:
version: 1.0.0
package:
pip: {name: reverse-words, version: 1.0.0} # index: for a private index
loader: load_environment # the default; or module:function
category: Custom
modes: [Train, Evaluate]
rewardType: Binary
```
## Run an environment from the Environments Hub
[Section titled “Run an environment from the Environments Hub”](#run-an-environment-from-the-environments-hub)
An environment published on [Prime Intellect’s Environments Hub](https://app.primeintellect.ai/dashboard/environments) is a pip package, so it runs on Nodus as it is. Name it with the Hub’s index for its owner and the version the Hub lists:
```yaml
apiVersion: nodus.dev/v1
kind: Environment
metadata:
name: reverse-text
spec:
version: 0.1.4
package:
pip:
name: reverse-text
version: 0.1.4
index: https://hub.primeintellect.ai/primeintellect/simple/
category: Custom
modes: [Train, Evaluate]
```
Its `load_environment()` returns a `verifiers` environment, and Nodus trains on it directly: the `dataset` is the train split, the `eval_dataset` the held-out one (with only one of them, every fifth task is held out), and each completion is scored by the environment’s own rubric, with `verifiers` 0.1 through 0.3. A multi-turn environment, such as a game like `wordle`, runs too: see [Train on a multi-turn environment](#train-on-a-multi-turn-environment), and so does one where the model calls tools: see [Train on a tool-calling environment](#train-on-a-tool-calling-environment). An environment whose tools run in a remote sandbox, a browser or an MCP server fails the build with that reason, as does one scored by another model (a `JudgeRubric`), since grading has no network and holds no credential for that model. The console marks each kind in its Hub search from the Hub’s own tags. The package comes from the Hub index alone and its dependencies may also come from PyPI.
Grading has no network. Nodus fetches the datasets an environment loads while it builds the image, and every task list and grade after that reads that copy, so a run always sees the data its build saw.
## Train on a multi-turn environment
[Section titled “Train on a multi-turn environment”](#train-on-a-multi-turn-environment)
In a multi-turn environment the model and the environment take turns: the model replies, the environment answers with feedback, and the episode goes on until the environment ends it. Set `maxTurns` on the TrainingJob (**Turns per episode** under **Advanced** in the console) to let an episode run that many model turns:
```yaml
spec:
runtime: nodus/grpo-lora
environment: {name: wordle@0.1.7, trainTasks: 512, heldOutTasks: 64, seed: 42}
parameters:
maxTurns: 6 # model turns per episode
maxCompletionLength: 256 # tokens per reply
maxEpisodeTokens: 2048 # tokens of the whole episode after the prompt, the model's and the environment's
```
The reward is the one the environment gives the whole episode, and training learns only from the model’s own tokens. An episode that reaches either cap ends there and is graded as it stands. With the default `maxTurns: 1`, only the first reply is graded. Baseline and final evaluation play the same episodes greedily.
An [OpenEnv](https://github.com/meta-pytorch/OpenEnv) environment runs in process from a package whose `load_environment()` returns it. Each task is a seed: Nodus resets the environment with it, and every action is the model’s reply as the one text field of the environment’s action class. The environment must play the same episode for the same seed. Seeds 0 to 799 train and 800 to 999 are held out, unless the environment sets `train_seeds` and `test_seeds`:
```python
import nltk
from textarena_env.server.environment import TextArenaEnvironment
def load_environment():
try: # the build downloads NLTK's word lists; grading reads them offline
nltk.data.find("corpora/words")
nltk.data.find("taggers/averaged_perceptron_tagger_eng")
cached = True
except LookupError:
cached = False
return TextArenaEnvironment("Wordle-v0", download_nltk=not cached)
```
A module of your own can also be multi-turn: give it `step(turns, answer)` in place of `reward`. It gets every model turn so far and returns the environment’s next message, or `None` once the episode ended, and the reward so far. Grading keeps no state between turns, so `step` replays the turns from the start.
## Train on a tool-calling environment
[Section titled “Train on a tool-calling environment”](#train-on-a-tool-calling-environment)
In a `verifiers` `ToolEnv` the model calls the environment’s tools and reads their results. Each task’s prompt offers the tools through the model’s chat template, the model calls one by writing a `{"name": "...", "arguments": {...}}` block (the form Qwen and most open chat templates teach), and the environment runs the call in its own code. The results are the next turn, and the episode ends when the model answers without calling a tool. Set `maxTurns` to the rounds of calls an episode may make:
```yaml
spec:
runtime: nodus/grpo-lora
environment: {name: tool-test@0.1.1, trainTasks: 21, heldOutTasks: 43, seed: 42}
parameters:
maxTurns: 3
maxCompletionLength: 256
```
The tools run inside the grader, which has no network, so a tool that searches the web or runs code in a remote sandbox cannot train. Pick a model whose chat template supports tools.
## Run any Reasoning Gym family
[Section titled “Run any Reasoning Gym family”](#run-any-reasoning-gym-family)
The catalog’s `nodus/reasoning-gym` serves five reviewed families. To train on any other [Reasoning Gym](https://github.com/open-thought/reasoning-gym) family, or a mix of them, publish a package whose `load_environment()` returns the dataset; its own `score_answer` scores each completion:
```python
import reasoning_gym
def load_environment():
return reasoning_gym.create_dataset("knights_knaves", size=2000, seed=7)
```
Every fifth entry is held out. The model’s last `…` is its answer, or the whole completion without one; partial credit is the reward and only a full score is `Correct`. Each family’s own licence applies.
## Run a verl or SkyRL dataset
[Section titled “Run a verl or SkyRL dataset”](#run-a-verl-or-skyrl-dataset)
A dataset prepared for [verl](https://github.com/volcengine/verl) or [SkyRL](https://github.com/NovaSky-AI/SkyRL) runs with its reward unchanged. Publish a module that names the parquet files with verl’s own keys, and your verl reward function under its usual name:
```python
train_files = "hf://datasets/BytedTsinghua-SIA/DAPO-Math-17k/data/dapo-math-17k.parquet"
val_files = "hf://datasets/BytedTsinghua-SIA/AIME-2024/data/aime-2024.parquet" # optional
def compute_score(data_source, solution_str, ground_truth, extra_info=None):
... # verl's custom reward signature; a dict's "score" also works
```
Each row is verl’s: `prompt` (a system message and one user message at most), `data_source`, `reward_model.ground_truth` and `extra_info`. Without `compute_score`, each row’s `env_class` names the SkyRL-gym environment that scores it, built from the row as SkyRL builds it; add `skyrl-gym` to the package’s dependencies. Without `val_files`, every fifth row is held out. The package needs `datasets` as a dependency, and a SkyRL environment of more than one turn fails the build with that reason.
To ship your own system packages or files, build the module into an image on the `env-base` image instead, push it and name its digest:
```dockerfile
FROM ghcr.io/nodus-compute/catalog/env-base:1.0.0
COPY reverse_words.py /opt/environment/
ENV NODUS_ENVIRONMENT_LOADER=reverse_words NODUS_ENVIRONMENT=reverse-words NODUS_ENVIRONMENT_VERSION=1.0.0
RUN nodus-env info
ENV HF_HUB_OFFLINE=1 HF_DATASETS_OFFLINE=1
```
```yaml
apiVersion: nodus.dev/v1
kind: Environment
metadata:
name: reverse-words
spec:
version: 1.0.0
package: {image: registry.example.com/acme/reverse-words@sha256:...}
category: Custom
modes: [Train, Evaluate]
rewardType: Binary
```
```console
$ nodus apply -f environment.yaml
$ nodus get environment/reverse-words -w # Ready once the image digest is verified
```
A TrainingJob names it without the `nodus/` prefix (`environment: {name: reverse-words@1.0.0}`), and in Python `rl.grpo_lora(environment="reverse-words@1.0.0", ...)`. The whole example, with a GRPO TrainingJob, is in [`examples/training/custom-reward`](https://github.com/nodus-compute/nodus-platform/tree/main/examples/training/custom-reward). Nodus pulls the image from a public registry or from your organization’s space in the Nodus registry; other private registries are not supported for Environments yet.
Any image works if it provides the two commands Nodus runs, as uid 10001 with no network:
* `nodus-env tasks --split train|test --seed N` writes one JSON line per task: `{"taskId", "prompt", "metadata"}`.
* `nodus-env grade` reads `{"taskId", "completion"}` lines and writes one `{"taskId", "verdict", "reward", "evidence"}` line for each, in order.
The Environment becomes `Ready` once its image is verified, with the declared split sizes in `status.splits`. The first TrainingJob that uses a split and seed runs `nodus-env tasks` in one of its own grading Sandboxes and Nodus keeps that manifest for every later run of your organization, so the tasks never change between runs. A version’s image and graders cannot change: publish a new version instead, so earlier results stay reproducible.
# Functions
> Run Python functions on Nodus workers, deploy them as an App that stays up, and look them up from anywhere.
A Function is a Python function that runs on Nodus workers. You decorate it, call it from your own code with `.remote()`, `.map()` or `.spawn()`, and Nodus starts workers when calls arrive, keeps them warm for a while and scales them back down. The pages of this guide cover [calling Functions](/docs/guides/functions/calls), [scaling them](/docs/guides/functions/scaling) and [what is billed while a worker waits](/docs/guides/functions/billing).
## Run an App
[Section titled “Run an App”](#run-an-app)
An App is the group of Functions in one file. `nodus run` creates an ephemeral App, runs your `main` on your machine and deletes the App when `main` returns. Every `.remote()` call runs on a Nodus worker.
examples/functions/map/app.py
```python
"""One Function called three ways, then mapped over a thousand inputs.
Run it with `nodus run examples/functions/map/app.py`.
"""
import nodus
app = nodus.App("fn-map")
@app.function(cpu=1, memory="1Gi", max_workers=4, target_concurrency=8, max_cost=1)
def square(x: int) -> int:
return x * x
@app.function(cpu=1, memory="1Gi", max_cost=1)
def divide(a: int, b: int) -> float:
return a / b
@app.local_entrypoint()
def main(n: int = 1000) -> None:
print("remote:", square.remote(7)) # one call; blocks for the result
call = square.spawn(8) # starts a call and returns a handle
print("spawned:", call.get(timeout=600))
results = list(square.map(range(n))) # one call per input, results in input order
print("map:", len(results), "ordered:", results == [i * i for i in range(n)])
try:
divide.remote(1, 0)
except ZeroDivisionError as exc: # the remote exception comes back as its own type
print("raised:", type(exc).__name__)
```
```console
$ nodus run examples/functions/map/app.py
remote: 49
spawned: 64
map: 1000 ordered: True
raised: ZeroDivisionError
```
The directory of the file (minus what `.gitignore` and `.nodusignore` exclude) is uploaded once per content hash, so the workers import the same code you ran. The App renews itself while `main` runs and is deleted shortly after `main` stops, even when your machine loses its connection.
## Deploy an App
[Section titled “Deploy an App”](#deploy-an-app)
`nodus deploy` keeps the App. Its Functions stay available after your program exits, and any program can call them.
examples/functions/warm-pool/app.py
```python
"""A Function that keeps one worker warm, so a call never waits for a start.
Deploy it with `nodus deploy examples/functions/warm-pool/app.py`. The idle worker is billed at the Function's
worker rate for as long as it stays warm; set `min_workers=0` to pay only while calls run.
"""
import nodus
app = nodus.App("fn-warm")
@app.function(cpu=1, memory="1Gi", min_workers=1, max_workers=3, scaledown_window="2m", max_cost=1)
def ping() -> str:
return "pong"
```
```console
$ nodus deploy examples/functions/warm-pool/app.py
```
```python
import nodus
ping = nodus.Function.from_name("fn-warm", "ping")
print(ping.remote()) # pong
```
A deploy updates the App in place. A Function whose code, image or resources changed rolls its workers once, after their in-flight calls finish. A Function you removed from the file is deleted, and the calls it still had queued end as `Failed` with the reason `FunctionDeleted`, so no caller waits for a call nobody will run.
## What Nodus creates
[Section titled “What Nodus creates”](#what-nodus-creates)
| Object | What it is | Look at it with |
| -------------- | ------------------------------------------------------------------------ | ------------------------------------------------------------ |
| `App` | The Functions of one file; ephemeral for `run`, persistent for `deploy` | `nodus get apps` |
| `Function` | One decorated function or class, with its image, resources and `scaling` | `nodus get functions` |
| `FunctionCall` | One invocation, kept for seven days after it ends | `nodus get functioncalls -l nodus.dev/function=fn-warm-ping` |
A Function reports its state in `status.phase`, how many workers it has in `status.workers`, how many calls wait in `status.queue`, and the expected start time in `status.estimate`.
```console
$ nodus get function fn-warm-ping
NAME PHASE WORKERS QUEUED COST AGE
fn-warm-ping Running 1/3 0 $0.02 3m
```
## Stop, restart and delete
[Section titled “Stop, restart and delete”](#stop-restart-and-delete)
```console
$ nodus stop function/fn-warm-ping # drain the workers; new calls wait in the queue
$ nodus start function/fn-warm-ping # start workers again for the calls that waited
$ nodus restart function/fn-warm-ping # replace every worker once, after in-flight calls finish
$ nodus delete app/fn-warm
```
A stopped Function keeps accepting calls and holds them in the queue, so work you submit while it is stopped runs once it starts. Workers also stop by themselves when your balance cannot cover another renewal; calls queue until you add credit and then run.
Note
A call that runs for longer than its `timeout` (five minutes by default, 24 hours at most) ends as `Failed` with the reason `DeadlineExceeded`, and its worker restarts, because the thread that ran it cannot be interrupted.
# What is billed while warm
> How Function workers are billed when they run calls, when they wait, and when they scale to zero.
You pay for worker time, by the second, at the worker’s rate. A call does not add a charge of its own. A Function shows the worker’s rate before you run anything, in `status.estimate.rateUSDPerHour` and in `f.estimate(...)`.
## When a worker is billed
[Section titled “When a worker is billed”](#when-a-worker-is-billed)
| State | Billed? |
| -------------------------------------------------------------------- | ----------------------------------- |
| A worker starting: placing, pulling its image, running `enter` hooks | Yes, from the moment it is placed. |
| A worker running at least one call | Yes. |
| An idle worker inside `min_workers` | Yes, always. This is warm capacity. |
| An idle worker above `min_workers`, inside `scaledown_window` | Yes, until the window ends. |
| A worker that was released | No. |
| A Function with no workers | No. |
So a Function with `min_workers=0` costs nothing while nothing is calling it, and each burst costs its workers’ time plus their `scaledown_window`. A Function with `min_workers=1` costs one worker’s rate every hour, and its calls never wait for a start.
## Reading the bill
[Section titled “Reading the bill”](#reading-the-bill)
Every worker is billed under one Function and shows up as two lines of usage: the time the worker spent running calls and the time it spent waiting (`warm_idle_seconds`). Both appear under the Function in your usage records, so the cost of keeping workers warm is never mixed into the calls themselves.
Each call’s `status.costUSD` is the worker time attributed to that call: the worker’s rate divided by `target_concurrency`, for the whole seconds the call spent on a worker, at least one. A worker that runs four calls at once shows each of them a quarter of the rate. Call costs show where the busy time went. The bill is the workers’.
## Spending limits
[Section titled “Spending limits”](#spending-limits)
A Function does not take a spend cap yet: setting `max_cost` is refused when you deploy, so no cap can be set and silently ignored. Your credit balance is the limit on what its workers spend, and an idle `min_workers` worker bills until you stop the Function.
When your balance cannot cover another renewal, the workers drain inside the reserve and the Function scales to zero. New calls queue and show `Funded=False` on the Function. Add credit and the Function starts workers again for the calls that waited. Nothing the calls were doing is lost: a call whose worker drained finishes first, and one that could not finish goes back to the queue.
# Call a Function
> Run one call, start calls without waiting, or map a Function over thousands of inputs, and read what comes back.
A Function has three ways to start a call. Each call is a `FunctionCall` object, so you can look it up later, wait for it from another process or cancel it.
| Call | What it does |
| ------------- | -------------------------------------------------------------------------------------------------- |
| `f.remote(x)` | Runs one call and blocks until it returns its result. |
| `f.spawn(x)` | Starts one call and returns a handle: `handle.get(timeout=...)` waits, `handle.cancel()` stops it. |
| `f.map(xs)` | Starts one call per input and yields the results in input order. |
```python
with app.run():
print(square.remote(7)) # 49
handle = square.spawn(8)
print(handle.get(timeout=600)) # 64
print(list(square.map(range(1000)))) # 1,000 results, in order
```
## Map over many inputs
[Section titled “Map over many inputs”](#map-over-many-inputs)
`.map()` creates its calls in batches of up to 1,000 per request, and a batch is all or nothing: a request that fails leaves no calls behind, and sending it again does not create them twice. Results come back in input order whichever worker finishes first. `order_outputs=False` yields them as they finish, and `return_exceptions=True` yields a failed input’s exception instead of stopping the loop.
```python
results = list(process.map(files, order_outputs=False, return_exceptions=True))
```
Several callers can map into one Function at the same time. Workers take calls from each caller in turn, so a caller that submits 10 inputs is not stuck behind another caller’s 10,000.
## Arguments and results
[Section titled “Arguments and results”](#arguments-and-results)
Arguments and results are serialized with cloudpickle. Values up to 64 KiB travel inside the call. Larger ones are uploaded once, by content hash, and the call carries a reference, so a large argument costs nothing extra to send again. The caller and the worker image must run the same Python minor version, 3.10 to 3.13.
## Errors, retries and timeouts
[Section titled “Errors, retries and timeouts”](#errors-retries-and-timeouts)
An exception raised inside the Function is raised again in your process as its own type when that type can be imported there. Otherwise you get `nodus.errors.RemoteError`. Either way the remote traceback is attached.
```python
try:
divide.remote(1, 0)
except ZeroDivisionError:
...
```
Two things can send a call back to the queue, and they are counted separately:
* **The function raised.** With `retries=nodus.Retries(max_retries=3)` the call is retried with exponential backoff, up to ten times. Without retries the exception ends the call.
* **The worker was lost.** A worker that disappears or is replaced does not count as a retry. Its calls go back to the queue and run on another worker, up to `recovery.maxAttempts` times (eight by default), then end as `Failed` with the reason `RecoveryLimitExceeded`.
A call can finish only once. A result that arrives from a worker that no longer holds the call is refused, so a call that was re-dispatched never ends with two results.
## Look a call up later
[Section titled “Look a call up later”](#look-a-call-up-later)
```python
call = nodus.FunctionCall.from_name("fn-warm-ping-bcdfghjklm")
print(call.get(timeout=60))
```
```console
$ nodus get functioncalls -l nodus.dev/map=