This is the full developer documentation for Nodus # Build with Nodus > Run your first job, choose a compute feature, and get from code to results. Run your code on cloud CPUs and GPUs. Start with a command, a Python function, or an interactive environment. > **From code to your first result** > > Install the CLI, sign in, and run a job in about five minutes. > > [Run your first job →](/docs/getting-started/) ## What are you building? [Section titled “What are you building?”](#what-are-you-building) Choose a starting point. Each guide covers the essentials and a working example. * **[Jobs](/docs/guides/jobs/)** Run a script or batch command to completion. * **[Workspaces](/docs/guides/workspaces/)** Develop with SSH, VS Code, or Jupyter. * **[Functions](/docs/guides/functions/)** Call Python remotely. Run calls in parallel. * **[Sandboxes](/docs/guides/sandboxes/)** Give agent-written code an isolated place to run. * **[Inference](/docs/guides/inference/)** Call a hosted model from your application. * **[Agents](/docs/guides/agents/)** Run agents with tools and saved progress. Looking for [training (Beta)](/docs/guides/training/), [storage](/docs/guides/volumes/), or [billing](/docs/guides/billing/)? [Browse all guides →](/docs/guides/) ## Bring your coding agent [Section titled “Bring your coding agent”](#bring-your-coding-agent) [Connect over MCP](/docs/for-agents/) to work from your editor, or give your agent the [Markdown task map](/docs/source/index.md). Read one relevant guide at a time; use the [reference](/docs/reference/) for exact commands and API fields. # For coding agents > Connect a coding agent over MCP, and the machine-readable versions of these docs, the API and the setup. Coding agents use Nodus the way you do: the same account, projects, budgets and confirmations. Connect one over MCP, or point it at the machine-readable outputs below. ## Read in this order [Section titled “Read in this order”](#read-in-this-order) 1. Read the [compute decision guide](/docs/source/guides/choose-compute.md) to choose the feature that fits the request. 2. If this is a first run, read the [quickstart](/docs/source/getting-started.md). 3. Use the [guide directory](/docs/source/guides.md) to find the relevant guide, then fetch its Markdown. 4. Look up exact commands in the [CLI reference](/docs/reference/cli/), signatures in the [Python reference](/docs/reference/python/), or request fields in [OpenAPI](/docs/openapi.json). Fetch individual pages to keep context small. For example, `/docs/guides/jobs/` has its Markdown at [`/docs/source/guides/jobs.md`](/docs/source/guides/jobs.md). Each docs page has a **View Markdown** link. Use the full corpus only for tasks that need many parts of the product. ## Connect over MCP [Section titled “Connect over MCP”](#connect-over-mcp) Nodus runs one MCP server with generic tools over every resource (`get`, `describe`, `logs`, `estimate`, `apply`, `exec` and more). Writes return a dry run first and run only once confirmed. | Client type | Configuration | | ---------------------------------------------------------- | -------------------------------------------------------------------------------------------- | | Hosted (Claude Code, Codex, Cursor and other HTTP clients) | Add the server URL from [`/mcp-hosted.json`](/mcp-hosted.json) and sign in with your browser | | Local stdio clients | Run `nodus mcp`; the configuration is [`/mcp.json`](/mcp.json) | Step-by-step setup for each client is at [/connect](/connect/), and as Markdown at [`/connect.md`](/connect.md). ## Machine-readable outputs [Section titled “Machine-readable outputs”](#machine-readable-outputs) | Output | What it holds | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------- | | [`/docs/pages.json`](/docs/pages.json) | Lightweight page manifest: titles, summaries, status and source URLs | | [`/docs/llms.txt`](/docs/llms.txt) | An index of these docs for language models | | [`Start`](/docs/bundles/start.md), [`compute`](/docs/bundles/compute.md), [`models`](/docs/bundles/models.md), [`data`](/docs/bundles/data.md), [`billing`](/docs/bundles/billing.md) | Focused bundles with complete pages, examples and source URLs | | [`/llms-full.txt`](/llms-full.txt) | Guides and concepts in one Markdown file; use the index for exact references | | `/docs/source/.md` | Each page’s Markdown, linked from the page with `rel="alternate"` | | [`/docs/index.json`](/docs/index.json) | Every page with its headings and text, versioned by build | | [`/docs/openapi.json`](/docs/openapi.json) | The HTTP API contract | | `/skills//SKILL.md` | Task instructions for agents that support skills | | [`/install`](/install), [`/install.ps1`](/install.ps1) | The CLI installer for macOS, Linux and Windows | A good first prompt for an agent that can read URLs: ```text Read https://nodus-compute.ai/connect.md and help me connect Nodus to this agent. Reuse any existing Nodus connection. Verify setup by listing my jobs. Do not start paid compute. ``` ## Use the current contract [Section titled “Use the current contract”](#use-the-current-contract) Fetch the relevant guide and its linked reference before writing commands. Keep the resource’s API version and Beta status explicit. Use the published OpenAPI schemas for field names and the CLI or SDK reference for the installed interface; do not invent flags from examples for another tool. Each Markdown page includes its canonical source URL and build revision. Cite the source page when explaining behavior. Read billing and recovery limits before creating work, and verify the resource status, logs and outputs before reporting success. A submitted request is not evidence that the run completed. # Get started > Install the Nodus CLI, sign in, run a command on a GPU, follow it, get its results and see exactly what it cost. This page takes you from nothing to a finished GPU job and its bill. You need a terminal and a browser. New accounts start with a **$30 starter grant**, so the first runs need no card. 1. **Install the CLI.** * pip ```sh pip install nodus-compute ``` The Python package (Python 3.10 or newer) includes the `nodus` CLI. * macOS and Linux ```sh curl -fsSL https://nodus-compute.ai/install | sh ``` * Windows ```powershell irm https://nodus-compute.ai/install.ps1 | iex ``` 2. **Sign in.** ```sh nodus login ``` Your browser opens to confirm the sign-in. A new account gets an org with a `default` project, an API key for this machine, and the $30 starter grant, which expires 30 days after it is granted. On a machine without a browser, run `nodus login --device` and approve the code from any other device. 3. **Run a command on a GPU.** ```console $ nodus run --gpu L4 --image nodus/pytorch -- python -c "import torch; print(torch.cuda.get_device_name())" job/run-4kq7z created · est. $0.01–0.03 · starts in ~2–4 min (cold) · Ctrl+C to cancel, -d to detach ✓ Scheduled l4-24g-x1-us · $0.52/h (rate frozen) ✓ Provisioning 1m48s ✓ Pulling image 22s ▶ Running NVIDIA L4 ✓ Succeeded in 2m31s · $0.02 · kept warm 60 s · nodus describe job/run-4kq7z ``` `nodus run` uploads the current directory, prints the estimate before anything is charged, streams the logs and exits with your command’s exit code. The rate is fixed when the machine is chosen and holds for the whole run. Add `-d` to return right away and let the job run on its own. 4. **Follow it.** Every run is a Job you can come back to. ```sh nodus get jobs -w # live list of your jobs nodus logs -f job/run-4kq7z # stream the logs again nodus describe job/run-4kq7z # status, events and the cost so far ``` 5. **Get the results.** Declare an output path when you run, then copy it back. Downloads are checked against their SHA-256 digest. ```sh nodus run --gpu L4 --image nodus/pytorch --output model=/outputs/model -- python train.py nodus cp job/:outputs/model ./model ``` 6. **See what it cost.** ```console $ nodus billing $ nodus get usage --field-selector object.name=run-4kq7z --group-by segment SEGMENT AMOUNT Boot $0.018778 Running $0.003033 Teardown $0.002889 ``` You pay for the machine from the moment it is created until it is deleted: starting up (`Boot`), your command (`Running`) and shutting down (`Teardown`), all at the rate frozen at launch. Credit is prepaid: a hold is reserved before a machine starts, and only what was used is charged. Using a coding agent? Connect Claude Code, Codex or Cursor to Nodus at [/connect](/connect/), then ask it to run your command. The agent uses the same account, projects and spending limits as the CLI. ## Next steps [Section titled “Next steps”](#next-steps) * [Concepts](/docs/concepts/): resources, projects, and how billing is measured. * [Pricing reference](/docs/reference/pricing/): list prices for every GPU, CPU shape, storage and egress. * [Guides](/docs/guides/): one guide per feature. # Find a guide > Choose a task, then read the guide you need. Start with running code and add storage, integrations, or billing as needed. New to Nodus? Start with [your first job](/docs/getting-started/). Otherwise, choose the task you have now. Not sure which feature fits? [Compare Jobs, Workspaces, Sandboxes and Functions](/docs/guides/choose-compute/). ## Run code [Section titled “Run code”](#run-code) | Task | Guide | | --------------------------------------- | --------------------------------------- | | Run a command to completion | [Jobs](/docs/guides/jobs/) | | Work in SSH, VS Code, or Jupyter | [Workspaces](/docs/guides/workspaces/) | | Run code in an isolated container | [Sandboxes](/docs/guides/sandboxes/) | | Call Python remotely or in parallel | [Functions](/docs/guides/functions/) | | Chain jobs together | [Pipelines](/docs/guides/pipelines/) | | Try multiple parameter combinations | [Sweeps](/docs/guides/sweeps/) | | Choose resources and check availability | [Compute selection](/docs/guides/gpus/) | ## Models and agents [Section titled “Models and agents”](#models-and-agents) | Task | Guide | | ---------------------------------------- | ------------------------------------------------------ | | Call a hosted model | [Inference](/docs/guides/inference/) | | Run agents with tools and saved progress | [Agents](/docs/guides/agents/) | | Fine-tune or train a model | [Training (Beta)](/docs/guides/training/) | | Train across machines | [Multi-node training (Beta)](/docs/guides/multi-node/) | | Define training and evaluation tasks | [Environments](/docs/guides/environments/) | ## Files and state [Section titled “Files and state”](#files-and-state) | Task | Guide | | -------------------------------------------- | ---------------------------------------- | | Package code and dependencies | [Images](/docs/guides/images/) | | Keep files between runs | [Volumes](/docs/guides/volumes/) | | Pass credentials to your code | [Secrets](/docs/guides/secrets/) | | Download results | [Outputs](/docs/guides/outputs/) | | Save and load progress after an interruption | [Checkpoints](/docs/guides/checkpoints/) | | Connect external data and services | [Connections](/docs/guides/connections/) | ## Tools and integrations [Section titled “Tools and integrations”](#tools-and-integrations) * [Python SDK](/docs/guides/python/): write and manage runs from Python. * [MCP](/docs/guides/mcp/): connect your coding agent to Nodus. * [Console](/docs/guides/console/) and [Ask Nodus](/docs/guides/assistant/): manage work in the browser. * [Logs and metrics](/docs/guides/logs/): follow progress and diagnose a run. * [Webhooks](/docs/guides/webhooks/) and [notifications](/docs/guides/notifications/): react to changes. ## Account, costs, and your own compute [Section titled “Account, costs, and your own compute”](#account-costs-and-your-own-compute) * [Account setup](/docs/guides/access/sign-up-and-orgs/), [projects](/docs/guides/access/projects/), and [members and roles](/docs/guides/access/members-and-roles/): organize your team. * [API keys](/docs/guides/access/api-keys-and-scopes/): give applications access. * [Billing](/docs/guides/billing/), [usage and costs](/docs/guides/billing/usage-and-costs/), and [budgets](/docs/guides/billing/budgets/): fund work and control spending. * [Pools](/docs/guides/pools/) and [cloud accounts](/docs/guides/pools/cloud-accounts/): use your own compute. ## Coming from another tool? [Section titled “Coming from another tool?”](#coming-from-another-tool) Start with [Nodus for Modal users](/docs/getting-started/modal-users/), [for kubectl users](/docs/getting-started/kubectl-users/), or [migrating from Nodus 0.x](/docs/guides/migrate-from-0x/). # How Nodus works > Resources, orgs and projects, how your work is placed and kept alive, and how prepaid billing measures it. This page is the map. Each section links to the concept page that goes deeper. ## Everything is a resource [Section titled “Everything is a resource”](#everything-is-a-resource) You describe work as a **resource**: a small declarative object with a `kind`, a `metadata.name` and a `spec`, the same shape Kubernetes uses. You create it with the CLI, the Python SDK, the console or the HTTP API, and Nodus reports progress in its `status`. The same object reads the same everywhere, so `nodus get`, `kubectl get` and the console show one truth. | You want to | Kind | Everyday command | | --------------------------------------------- | -------------------------------------------------------- | ---------------------------------------- | | Run a command to completion | `Job` (and `Pipeline`, `Sweep` to chain or fan out Jobs) | `nodus run`, `nodus apply -f job.yaml` | | Keep an isolated container for untrusted code | `Sandbox` | `nodus create sandbox`, `nodus exec -it` | | Call Python remotely and fan out | `App`, `Function`, `FunctionCall` | `nodus deploy app.py` | | Run a durable agent | `Agent`, `AgentRun`, `AgentGroup` | `nodus create agentrun` | | Serve or call a model | `Model`, `InferenceEndpoint` | `nodus get models -n nodus` | | Train or evaluate with a recipe | `TrainingJob`, `TrainingRuntime`, `Environment` | `nodus create trainingjob` | | Develop on a remote machine | `Workspace` | `nodus ssh workspace/` | | Store data, images and secrets | `Volume`, `Image`, `Secret`, `Connection` | `nodus volume put` | | Bring your own machines | `Pool`, `Node`, `EnrollmentToken`, `CloudAccount` | `nodus create pool` | `TrainingJob` and `TrainingRuntime` are Beta. Every other kind above is generally available. Run `nodus api-resources` for the full list and `nodus explain .spec` for any field. ## Orgs and projects [Section titled “Orgs and projects”](#orgs-and-projects) An **org** is your billing and access boundary: members, API keys, credit and budgets belong to it. Inside an org, **projects** group resources (every org starts with `default`). Pass `-p ` to the CLI, or set `NODUS_PROJECT` for a whole shell. Names are unique within a project, and labels such as `team=nlp` let you select resources and break down cost across projects. ## How your work runs [Section titled “How your work runs”](#how-your-work-runs) You state requirements (an accelerator and count, memory, a region class, a deadline or a cost ceiling) and Nodus chooses an **offering** that satisfies them, such as `h100-sxm-80g-x8-us`. You see offerings and Nodus ids, never the machines behind them. Each placement of your container on a machine is an **Attempt**. When capacity is reclaimed, Nodus starts a new Attempt and restores the files your program saved to its checkpoint directory (`NODUS_CHECKPOINT_DIR`). Your program reloads its own model, optimizer and progress from those files; Nodus restores files, not process memory. ## How billing works [Section titled “How billing works”](#how-billing-works) Nodus sells **prepaid credit**, and four rules decide every charge: 1. **A hold comes first.** Before any paid machine starts, a hold reserves enough credit on your org, and on every budget that applies, to cover the expected run. A launch that cannot be funded is refused with the amounts and a fix, and a running Job stops gracefully, inside its reserved amount, when money runs out. 2. **The rate is frozen at launch.** `nodus run` and the console show the estimate before launch; the rate chosen when the machine is acquired is the rate for the whole run, and it is never above the published list price. 3. **You pay what the provider bills for your machine, and nothing Nodus caused.** Rented capacity bills from the moment the machine is created until its deletion is confirmed, at provider cost ÷ 0.875 (Nodus keeps 12.5 % of what you pay). Capacity Nodus chose and discarded, orphaned machines and Nodus failures are never charged. 4. **Usage is itemized by segment.** Every compute usage line names the part of the machine’s billed time it covers: | Segment | Covers | | ---------- | ---------------------------------------------------------------------------------------- | | `Boot` | From the machine’s creation to your command starting: start-up, readiness and image pull | | `Running` | Your command, until it stops | | `Restore` | A replacement machine’s start-up when your work resumes from a checkpoint | | `Teardown` | From stop to confirmed deletion, plus the provider’s rounding increment | `nodus get usage --group-by segment` and the console’s Cost tab show the breakdown for any object. Nodus-operated capacity (Sandbox nodes, warm pools, CPU nodes) bills from the pricebook’s published per-vCPU, per-GiB and per-disk rates instead. Storage above 10 GB per org and egress above 10 GiB per org per day are metered; logs are free. The starter grant The first org a verified user creates receives **$30 of credit that expires 30 days after it is granted**. Grant credit is spent before purchased credit, soonest-expiring first. Until your first purchase, the org has starter limits (one Nodus node, three live Sandboxes, Sandbox lifetimes up to two hours), which lift when you buy credit. See the [pricing reference](/docs/reference/pricing/) for every published rate. # Attempts and recovery > How a run survives a reclaimed or lost machine, what each continuity mode keeps, and who pays for the time a recovery takes. Every run on Nodus executes as one or more **attempts**. An attempt is one incarnation of your command on one machine. When that machine is reclaimed, loses its network or fails, Nodus starts a new attempt somewhere else and your run continues from what it saved. This page explains what carries over, how Nodus decides a machine is gone, and what you pay for along the way. ## Attempts and epochs [Section titled “Attempts and epochs”](#attempts-and-epochs) A Job index, a Sandbox or a worker slot holds **one running attempt at a time**. Each new attempt gets the next **epoch**, a number that only grows. Nodus accepts reports, checkpoints and outputs only from the current epoch, so a machine that comes back after it was replaced can never overwrite the work of its successor. You see a run’s attempts with `nodus get attempts -l nodus.dev/job=`. A run’s attempts share its name with an epoch suffix, and gang members add a rank suffix (`-r1`, `-r2`). | Attempt phase | Meaning | | --------------------------------- | ------------------------------------------------------------------------------------------------- | | `Pending`, `Placing`, `Acquiring` | Choosing and preparing capacity | | `Starting` | The machine is ready; the image is pulled and inputs or a checkpoint are restored | | `Running` | Your command is running | | `Succeeded` | Your command exited 0 | | `Failed` | Your command or its machine failed; `reason` says which (`NodeLost`, `Preempted`, `OOMKilled`, …) | | `Cancelled` | The attempt was stopped on purpose: a suspend, a cancel, a budget or lifetime limit | ## Continuity modes [Section titled “Continuity modes”](#continuity-modes) `recovery.continuity` says what a new attempt starts from: | Mode | A new attempt starts with | Use it for | | -------------- | --------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------ | | `Checkpointed` | The latest committed checkpoint of your declared state paths (`NODUS_CHECKPOINT_DIR` by default) | Training and long jobs that save their own model, optimizer and progress files | | `Restartable` | A cold start plus the progress cursor you reported (`NODUS_CURSOR_COMPLETED`, `NODUS_CURSOR_TOTAL`) | Batch work that can skip what it already finished | | `Ephemeral` | A cold start | Short or idempotent work | | `Snapshotted` | The latest filesystem snapshot | Sandboxes | Restoring files restores **files only**, never process memory. Your program loads its own checkpoint when it starts; Nodus never adds resume flags to your command. An empty checkpoint never counts as saved progress and never replaces an earlier useful one. If every checkpoint a run commits is empty, the run shows `Checkpointed=False, reason=NotCheckpointable`, and Nodus plans and prices it as `Ephemeral` from then on. ## What survives a preemption [Section titled “What survives a preemption”](#what-survives-a-preemption) When a provider reclaims interruptible capacity or a machine stops answering, Nodus recovers **make-before-break**: 1. On a reclaim notice, Nodus asks your attempt for an **urgent checkpoint** and, at the same time, starts preparing a replacement machine. 2. The replacement is prepared up to the point where it could start, but it does **not** start while the old machine might still be writing. 3. It starts only once the old machine is provably gone: the old attempt acknowledged its stop, the provider confirmed the machine terminated, or the old attempt’s lease ran out. 4. The new attempt restores according to the continuity mode above. So a `Checkpointed` run loses at most the work since its last committed checkpoint, and a checkpoint committed during the reclaim notice still counts. If the old machine comes back before it is replaced, the replacement is released and your run simply continues; you are not charged for the replacement. A machine that stops sending heartbeats gets a replacement prepared after 20 seconds. After 60 seconds of silence an attempt on shared capacity is declared lost and replaced. An attempt on a machine dedicated to your organization may keep running up to the edge of its funding while its replacement waits, because such a machine can only lose its own work. ## Recovery limits [Section titled “Recovery limits”](#recovery-limits) Recovery stops, and the run fails with a typed reason, when: | Reason | Rule | | ----------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `NoProgress` | Two attempts in a row made no progress. Progress means the command started and either ran for `recovery.minProgressDuration` (default 2 minutes) or committed a non-empty checkpoint. A preemption after progress never counts | | `RestoreFailed` | The same checkpoint failed to restore twice | | `ImagePullFailed` | The image failed to pull on a second machine (a pull failure is retried once elsewhere) | | `MaxAttemptsExceeded` | `recovery.maxAttempts` recoveries were used (default 8; 3 for distributed Jobs) | | `Preempted`, `NodeLost` | `recovery.onInterruption: Fail` was set, so the first interruption ends the run | Failures caused by your command (a non-zero exit, out of memory, an invalid checkpoint) fail fast and are not retried. At most three recoveries per organization prepare capacity at the same time; the others wait their turn and show an Event. ## Suspending [Section titled “Suspending”](#suspending) `nodus suspend job/` stops the run after a final checkpoint and releases its compute; `nodus resume` continues at a new epoch from that checkpoint. If the final checkpoint fails or takes longer than `max(10 minutes, 2 × the shutdown reserve)`, the run **keeps running** with `Suspended=False, reason=SuspendFailed` and a `SuspendFailed` Event, so a suspend never throws work away. Stops caused by money (credits, budgets, the maximum cost) or by a lifetime limit do not wait: they stop at the funded edge with whatever checkpoint exists. A stop always wins over recovery. If the machine is lost while a suspend, cancel or money stop is in progress, the attempt ends `Cancelled` with the stop’s reason and nothing is restarted or billed again; `nodus resume` starts the next epoch as usual. ## Who pays for recovery [Section titled “Who pays for recovery”](#who-pays-for-recovery) You pay for the machines your run uses, at the rate frozen when each was acquired, from the moment the provider starts billing through the confirmed deletion of the machine. The bill splits each machine’s time into segments: | Segment | Covers | | ---------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `Boot` | Start-up: boot, image pull, readiness, and for gangs the wait at the start barrier | | `Restore` | Start-up of an attempt that resumes: a recovery, a resume after a suspend, and for gangs the surviving machines’ wait from the restart to the next epoch’s start | | `Running` | Your command running | | `Teardown` | Stop to confirmed deletion, plus the provider’s billing increment | A recovery therefore costs you the replacement’s `Boot` or `Restore` time and the lost machine’s `Teardown`. A machine the provider refuses, or that never appears, before it is ready is replaced with another one, up to six tries per attempt; after that the attempt fails with `ReadinessFailed` and the run’s retry rule applies. Nodus pays, and never charges you, for: * a replacement released because the old machine came back; * extra machines prepared to start faster (hedges) that did not win; * machines that failed Nodus’s own checks after creation, and Nodus internal failures; * for distributed Jobs, probe failures on a qualified network path and outages of the Nodus mesh. `nodus billing usage --group-by segment` and the Cost tab show every segment of every attempt. # How billing works > Prepaid credits, holds before every paid action, exact per-second charges, and what happens when money runs out. Nodus is prepaid. You add credits, every paid action reserves funds before it starts, and charges come from the rate frozen when the capacity was acquired. Work never runs on credit you do not have, and it stops gracefully, with its progress saved, when money runs out. ## Your balance [Section titled “Your balance”](#your-balance) Your org has one balance, made of buckets: * **Purchased credit** from top-ups you pay for by card. * **Credit grants**: the starter credit, promo codes and credits from Nodus support. Each grant is its own bucket and may expire. Charges draw from grants first, soonest expiry first, then from grants that never expire, then from purchased credit. That way a grant is used before it expires, and purchased credit, the only kind that can be refunded, lasts longest. `nodus billing` and **Usage & billing → Overview** show: | Field | Meaning | | ------------------ | ----------------------------------------------------------------- | | Available | What you can spend now: your buckets minus open holds | | Reserved | Open holds, with the objects that hold them | | Purchased, credits | What is left in each kind of bucket, and the next grant to expire | | Arrears | Unpaid charges; new work waits until a top-up settles them | ## Holds and captures [Section titled “Holds and captures”](#holds-and-captures) Before a Job, Sandbox, Workspace, Function worker, agent run or build starts, Nodus places a **hold**: enough funds for the first stretch of work plus the cost of stopping it cleanly. While the work runs, Nodus captures what it used every 5 minutes and renews the hold for the next stretch. When the work ends, the final capture charges the exact amount and the rest of the hold is released. The estimate before launch shows the hold, the expected cost range and the minimum charge. A create that cannot be funded fails at once with the amounts, instead of queuing and failing later: ```text Error from server (InsufficientCredits): job "train-a" needs a $3.20 hold to start (released when it ends); available $1.10. fix: nodus billing top-up 20, or lower spec.maxCostUSD ``` ## What is charged [Section titled “What is charged”](#what-is-charged) Machines Nodus acquires for you are charged per second from the moment the provider starts billing until the machine is confirmed deleted, at the rate frozen when it was acquired. Usage itemizes each machine’s time by segment (`Boot`, `Restore`, `Running`, `Teardown`). [What you pay for](/docs/guides/billing/what-you-pay-for/) lists every kind of time and who pays for it, and [Pricing](/docs/concepts/pricing/) explains how rates are set. ## Limits [Section titled “Limits”](#limits) Three limits apply to every hold, and the tightest wins: * **Your balance.** A hold never exceeds what is available. * **Budgets.** A `Block` Budget over the org, a project or a label selector stops new holds and renewals in its scope when it is exhausted. * **Object caps.** `spec.maxCostUSD` on an object bounds that object and everything it owns. When a limit is reached, running work stops gracefully: Jobs suspend after a checkpoint, Sandboxes and Workspaces stop, and agent runs wait. Each stopped object shows `Funded=False` with the reason, and resumes when funds return. If a charge would still go past the limit, Nodus absorbs the difference. ## Low balance [Section titled “Low balance”](#low-balance) You get a low-balance warning by email, webhook and a console banner when your available balance falls below $5.00 (configurable) or below what your open holds need for their next renewal. Turn on [auto-recharge](/docs/guides/billing/#auto-recharge) to top up automatically below a threshold. ## Arrears [Section titled “Arrears”](#arrears) A charge that arrives after its hold is gone, such as daily storage, can leave unpaid charges. While arrears are open, new holds and uploads are refused with `402 ArrearsOutstanding`; your next top-up or grant settles them first. # The resource model > How every Nodus object is named, stored, changed, watched and deleted, and how that maps to Kubernetes tools. Everything you run on Nodus is an **object**: a Job, a Sandbox, a Function, a Volume, a Budget. Every object has the same shape and the same verbs, so once you know one kind you know them all, and the CLI, the Python SDK, the console, MCP clients and `kubectl` all work the same way. ## The shape of an object [Section titled “The shape of an object”](#the-shape-of-an-object) ```yaml apiVersion: nodus.dev/v1 kind: Job metadata: name: train-a # unique per project and kind namespace: default # the project labels: app: trainer spec: # what you want image: nodus/pytorch:2.8-cuda12.8 command: [python, train.py] maxCostUSD: "40.00" status: # what Nodus observed; written by Nodus only phase: Running ``` * **`metadata.name`** is a DNS label of at most 63 characters, unique within its project and kind. Set `metadata.generateName` instead to get a unique suffix; such a create needs an `Idempotency-Key` header. * **`metadata.namespace`** is the project. Every org starts with the project `default`. The read-only project `nodus` holds what Nodus publishes, such as base images, which you reference as `nodus/:`. * **`metadata.uid`** is a stable id such as `job_01j9…`, and **`metadata.resourceVersion`** changes on every write. Send the `resourceVersion` you read with an update to make it conditional: if someone changed the object in between, the update fails with error code `Conflict` and you read it again. * **Labels** select objects (`-l app=trainer`). Labels and annotations under the `nodus.dev/` prefix are set by Nodus; the label `nodus.dev/created-by` records who created an object and `nodus.dev/launched-by` from where. ## Desired state [Section titled “Desired state”](#desired-state) You change what an object does by changing its `spec`. Stopping, suspending and cancelling are values of `spec.state`, not separate actions: `nodus suspend job/train-a` sets `spec.state: Suspended`, and a later `nodus apply` of the same file keeps it suspended. Actions that change nothing in `spec`, such as restarting a Sandbox’s service, are requests: `nodus request restart sandbox/dev` sets the annotation `nodus.dev/restart-requested-at` to the current time, and Nodus acts once per value it has not handled yet. Most `spec` fields cannot change after create. An update that changes one fails with error code `FieldImmutable`, and the response lists each field. The fields that can change, such as `spec.state` and a raise of `spec.maxCostUSD`, are listed in each kind’s reference. `status.phase` uses the same words on every kind, so `--field-selector status.phase=Running` means the same thing everywhere. Run-to-completion kinds move through `Queued`, `Provisioning`, `Running` and end in `Succeeded`, `Failed` or `Cancelled`; long-running kinds move through `Pending`, `Starting`, `Running` and `Stopped`. `status.conditions` explain the details, such as the condition `Ready`. ## Create, review, apply [Section titled “Create, review, apply”](#create-review-apply) * **Create by name is safe to retry.** Creating an object that already exists with the same spec returns the existing object. A different spec fails with error code `AlreadyExists` and lists the differences. * **Review before you pay.** Add `?dryRun=All` (`--dry-run=server` in the CLI) to run every check without creating anything. The response shows the object with its defaults and, for kinds that use compute, `status.estimate`: the expected cost, the first hold and the start time. Send the `ETag` of that response as `If-Match` on the real create to launch exactly what you reviewed: if the spec or the price book changed, the create fails with error code `PreconditionFailed` and you review again. If only the hourly rate moved, the create goes ahead and never runs above the reviewed rate plus 10 % (unless you set `placement.maxRateUSDPerHour`); when nothing fits, the object waits in `Queued` and its status says the price is above the estimate. * **Retries never double-create.** The CLI and SDK send an `Idempotency-Key` with every create. Retrying with the same key returns the first response with the header `Idempotent-Replayed: true`; reusing a key for a different request fails with error code `IdempotencyKeyReused`. ## Watch instead of polling [Section titled “Watch instead of polling”](#watch-instead-of-polling) Every list can be watched: `nodus get jobs -w`, or `?watch=true&resourceVersion=` on the API. A watch delivers every change after that point, in order, and never skips one. Watches send newline-separated events (`ADDED`, `MODIFIED`, `DELETED`) as JSON, or as `application/x-ndjson` when you ask for it; with `allowWatchBookmarks=true` you also get a bookmark every 30 seconds to resume from. Nodus keeps 24 hours of changes: resuming from an older point fails with error code `Expired`, and you list again. Lists return at most 500 objects by default (`limit`, up to 1,000) and a `continue` token for the next page. ## Deleting [Section titled “Deleting”](#deleting) `nodus delete job/train-a` stops the work first and removes the object once cleanup is done: while it waits the object shows a `deletionTimestamp`, and the phase `Cancelling` or `Terminating`. Nodus tracks that cleanup with finalizers such as the finalizer `nodus.dev/billing`, which clears when the final charge has posted, so you never see an object gone while it is still costing money. Deleting an object also deletes what it created (a Pipeline’s Jobs, for example); with `--cascade=orphan` those stay. ## Errors [Section titled “Errors”](#errors) Every error has the same body: a Kubernetes `Status` whose `reason` is a stable error code, plus `fix` (what to do next), `docs` (a page for that code) and `requestId` (quote it to support). The [error reference](/docs/reference/errors/) lists every code. ## Kubernetes tools [Section titled “Kubernetes tools”](#kubernetes-tools) The API speaks the Kubernetes wire protocol for these objects, so `kubectl`, `k9s` and client-go work against it: projects are namespaces, `kubectl get jobs.nodus.dev` lists Jobs, and `kubectl explain`, `apply`, `diff`, `wait` and `-w` behave as they do on a cluster. The `nodus` CLI adds what `kubectl` does not have, such as logs, exec and file transfer for Nodus kinds. # Durable execution > How an agent run survives restarts, and what it needs from your code to do so. An agent run is recorded as it goes. Each effect it has is a **step**: creating its sandbox, choosing a model for a prompt, each model call and each command in the sandbox. Nodus stores the result of every step the first time it runs. If Nodus restarts, or a machine fails, the run is picked up again and replays from the top: every step that was recorded returns its stored result, and the run continues with the first step that was not. ## What this guarantees [Section titled “What this guarantees”](#what-this-guarantees) * A recorded step never runs again for the same run: no second model call, no second command, no second charge. * Pure and idempotent steps can retry after an interruption. Nodus model calls reuse their inference key. * An interrupted sandbox command or model call using your own key can have an uncertain outcome. Its run parks in `Waiting` with reason `NeedsResolution` instead of repeating the effect. * A run is picked up again by one dispatcher at a time. If two start, one stops without writing. * A run that reaches a different step than the one recorded at the same position fails with `NonDeterministicReplay`, and its record stays readable. ## What survives a restart [Section titled “What survives a restart”](#what-survives-a-restart) The run’s conversation, its answer so far and its status survive, because they come from the recorded steps. The sandbox’s files are not part of the record. If a sandbox is lost, the run creates a new one from the agent’s image and the files are not restored, so keep what must outlive a sandbox in the run’s answer. ## The determinism contract [Section titled “The determinism contract”](#the-determinism-contract) Code that runs between steps must give the same result for the same inputs and step results. Clock reads, random numbers and network calls belong inside steps. Agents on Nodus’s Claude follow this for you: the loop that runs the model and its commands is Nodus’s. ## Where the record lives [Section titled “Where the record lives”](#where-the-record-lives) Step results are encrypted with your organisation’s key. A run’s recorded payloads expire 30 days after it ends; its status, its cost and its step list stay. ## When an external outcome is uncertain [Section titled “When an external outcome is uncertain”](#when-an-external-outcome-is-uncertain) Inspect the run’s steps and the affected files or external system before starting replacement work. A `Started` step means the intent was saved but its outcome was not recorded; it does not prove the command failed or succeeded. Sending another message does not resume this run. Cancel it with `nodus cancel agentrun/` to release its sandbox. The sandbox can continue billing while the run waits. A configured deadline still ends the run. The managed step-resolution API is not available yet. Do not treat a new run as a safe retry until you have checked the earlier effect. # Gang networking > How the members of a distributed Job reach each other across machines and providers, how Nodus checks the path before training starts, and what each path can carry. Beta Distributed Jobs (`Job.spec.distributed`) are in beta behind per-org access. Gangs have 2 to 8 members, and `Relayed` gangs have 2 until multi-member Relayed runs are qualified. Ask for access from the console. A distributed Job runs as a **gang**: one member per machine, every member started together, sometimes on machines from different providers. Before your command starts, Nodus joins the members into a private network that belongs to that gang alone. Each rank gets the addresses of the others in its environment, so `torchrun`, Ray and plain `torch.distributed` work without any networking code of yours. You choose two things in `spec.distributed`: * `network`: how close the members must be. `Colocated` keeps them in one provider region, `Regional` allows any provider inside one region class, and `Global` allows anywhere. * `transport`: which paths you accept. `Direct` (the default) accepts only `Private` and `Direct` paths. `Auto` also accepts `Relayed` paths, which lets members on machines without a kernel network device join. Nodus picks the path. Every member of a gang uses the same path class, so every rank sees the same addressing and the same NCCL settings. ## Path classes [Section titled “Path classes”](#path-classes) | Path | When Nodus uses it | How traffic flows | Encryption | Bandwidth | | --------- | ----------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------- | ------------------------------------------------ | ------------------------------------------------------------------- | | `Private` | Every member is in one provider region with a private network | The provider’s private network, opened only between the gang’s members | None added (the provider’s network) | Provider native | | `Direct` | Every member can create a WireGuard device and the probe finds a direct path between every pair | WireGuard between the members, on a network device named `nodus0` | WireGuard | Measured per pair (see [measured throughput](#measured-throughput)) | | `Relayed` | A member has no network device of its own (container-only offerings) and `transport: Auto` | WireGuard in user space, carried through Nodus relays | WireGuard end to end; relays see only ciphertext | **Low: small models only** | ### Private [Section titled “Private”](#private) Members in the same provider region talk over the provider’s private network. Nodus opens the rendezvous and NCCL ports between the gang’s members only and closes them when the gang ends. The advertised addresses are the members’ private IPs. ### Direct [Section titled “Direct”](#direct) Each member runs a WireGuard endpoint next to your container, never inside it: your container gets no extra privileges. The members find each other through the Nodus mesh control and connect peer to peer, punching through NAT where needed. The advertised addresses are the members’ mesh addresses (`100.64.0.0/10`), and NCCL and Gloo use the `nodus0` interface. `Direct` means direct. If the probe finds that one pair can only connect through a relay, a `transport: Direct` gang is placed again without that pair, and a `transport: Auto` gang continues as `Relayed`. If a running pair loses its direct path, its traffic falls back to a relay without interrupting training, and the Job’s `NetworkDegraded` condition becomes true. ### Relayed [Section titled “Relayed”](#relayed) Container-only offerings give your container no network device, no `CAP_NET_ADMIN` and no UDP. Their members can still join a gang: each runs WireGuard in user space and sends peer traffic through two relays that Nodus operates. To reach that WireGuard endpoint, your processes start with a small preloaded library, `libnodus-netshim.so`, which redirects only connections to the other members of your gang. Everything else, such as downloads, datasets, object storage and model APIs, goes out directly as usual. The advertised addresses are the members’ own `eth0` addresses, so `MASTER_ADDR`, the `torchrun` rendezvous and the listeners NCCL opens on ephemeral ports all work unchanged, in both directions. The relays mesh with each other: if one goes down, members move to the other within seconds and the TCP connections inside WireGuard survive. ## What every rank sees [Section titled “What every rank sees”](#what-every-rank-sees) The path decides a few variables; the rest of the distributed environment is the same on every path. | Variable | `Private` | `Direct` | `Relayed` | | ---------------------------------------------------------- | --------------------- | -------------- | -------------------------------------------------- | | `NODUS_GANG_TRANSPORT` | `private` | `direct` | `relayed` | | `MASTER_ADDR`, `PET_RDZV_ENDPOINT`, `NODUS_NODE_IPS` | Private IPs | Mesh addresses | `eth0` addresses | | `NCCL_SOCKET_IFNAME`, `GLOO_SOCKET_IFNAME` | The private interface | `nodus0` | `eth0` | | `LD_PRELOAD`, `NODUS_NETSHIM_PEERS`, `NODUS_NETSHIM_SOCKS` | Not set | Not set | The shim first, then your image’s own `LD_PRELOAD` | Your own `NCCL_*` values, such as `NCCL_DEBUG=INFO`, win, except `NCCL_NET`, `NCCL_SOCKET_IFNAME`, `NCCL_SOCKET_FAMILY` and `NCCL_IB_DISABLE`, which the path fixes. A spec `LD_PRELOAD` is rejected when `transport: Auto`, because Nodus needs to place the shim first on a `Relayed` path. An `LD_PRELOAD` set in the image, such as jemalloc or tcmalloc, is kept after the shim. ## One network per epoch [Section titled “One network per epoch”](#one-network-per-epoch) A gang’s network lives exactly as long as one **epoch** of the gang: one attempt at running it with a fixed set of members. Each epoch gets a fresh mesh identity with single-use join keys that expire after 15 minutes, and its members can reach only each other. When Nodus restarts the gang, for example to replace a lost member, it fences the old epoch first: the old network is torn down, and a member left behind on it can no longer reach the new gang, nor can the new gang reach it. Members of other gangs, in your org or anyone else’s, are never reachable. ## The probe [Section titled “The probe”](#the-probe) Before your command starts on any member, every member checks its path to every other member. Nodus bills this time as boot time. 1. **Readiness.** Each member opens three fresh TCP connections to every peer, on ephemeral ports, within 60 seconds, through the same path your processes will use. 2. **Measurement.** For each pair, the member records the round-trip time, a 10-second TCP throughput sample and the path class it actually got: `direct`, or `derp:` when the pair goes through a relay. 3. **Shim self-test** (`Relayed` only). The member runs a check program with your image’s own loader to confirm that the shim loads. The slowest pair is published in `status.gang.probe` as `tcpGbps` and `rttMs`, with its path, and sets the `NetworkQualified` condition. A pair that cannot connect, or a path worse than your `transport` allows, fails the epoch and Nodus places the gang again; that time is not billed to you. The one exception is an image whose programs cannot load the shim: placing the gang elsewhere cannot fix it, so the Job fails at once with `ShimNotLoaded` and the time is billed (see [caveats](#caveats)). `tcpGbps` is a TCP throughput sample between two members, not NCCL bus bandwidth. Read it as an upper bound on what one connection between that pair can carry. ```console $ nodus get job llama-ft -o jsonpath='{.status.gang.probe}' {"tcpGbps":0.41,"rttMs":38.2,"path":"derp:nodus-us","measuredTime":"…"} ``` ## Measured throughput [Section titled “Measured throughput”](#measured-throughput) Nodus publishes only measured numbers: the probe’s TCP throughput and RTT, and NCCL all-reduce bus bandwidth from a separate benchmark on 2 × 1 H100. Nodus places a gang only on combinations of offerings it has qualified for the path, and a row appears here once its qualification run passes. | Path | Members | TCP throughput (probe) | RTT | NCCL bus bandwidth | | --------- | ----------------------------------- | ---------------------- | ----------------- | ------------------ | | `Private` | One provider region | Not yet published | Not yet published | Not yet published | | `Direct` | VMs in different providers | Not yet published | Not yet published | Not yet published | | `Relayed` | A container-only offering with a VM | Not yet published | Not yet published | Not yet published | Until the benchmark publishes, treat `Relayed` as low bandwidth: the scheduler assumes it is four times slower than `Direct` when it compares placements, and the estimate warns about it. ## Billing [Section titled “Billing”](#billing) `Private` and `Direct` traffic costs nothing beyond the members’ own time. `Relayed` traffic crosses Nodus relays and is billed per GiB on the `mesh-relay` line of your usage, measured on the relays as the bytes each member sends. Mesh traffic never counts as container egress. Each `Relayed` gang may send up to 1 Gbit/s through the relays, split evenly across its members; the relays are shared and best effort within that limit. ## Caveats [Section titled “Caveats”](#caveats) * **Synchronous training across providers is bound by the WAN.** Expect cross-provider data-parallel training to be limited by network bandwidth and latency, not by the GPUs. `Relayed` suits small models, algorithms that communicate little, reinforcement learning with separate rollouts, and getting N GPUs now wherever they are. * **`Relayed` needs dynamically linked glibc programs.** The shim is a preloaded library, so it loads only into programs that use the system’s glibc loader. Launchers built on musl (Alpine images) or linked statically never load it and cannot reach their peers. The probe detects this and the Job fails with `ShimNotLoaded`. Fix it by using a glibc-based image (any Debian, Ubuntu or CUDA image works), or by setting `transport: Direct`. * **`Relayed` is low bandwidth.** All gang traffic crosses the relays, under the per-gang limit above. Nodus makes no throughput claim for `Relayed` until the benchmark is published. * **Advertised addresses must be distinct and exclusive.** On `Relayed`, each member is reached at its own `eth0` address, so no two members of a gang may share one, and no two running `Relayed` gangs may route the same address. When placement produces such a collision, Nodus places the members again; it shows as `IPCollision` in the `GangReady` condition while it does. * **Only connections to peers go through the mesh.** On `Relayed`, connections to other members are redirected; UDP and connections to any other address are not. Tools that need UDP between members, or connect to peers through a hostname that does not resolve to their advertised address, do not work on `Relayed`. * **The probe measures TCP.** A good `tcpGbps` does not guarantee NCCL performance; NCCL’s numbers come only from the benchmark in the table above. # Lifecycles > The phases a Job, Pipeline or Sweep moves through, what moves it, and how spec.state and conditions relate. Jobs, Pipelines and Sweeps share one lifecycle: they run until they finish. Each shows where it is in `status.phase`, why in `status.reason` and `status.message`, and the details in `status.conditions`. You steer it with `spec.state`. ## Phases [Section titled “Phases”](#phases) | Phase | Meaning | Billed | | -------------- | ----------------------------------------------------------------------------------- | ------------------------------- | | `Queued` | Admitted; waiting for capacity that fits and for a funded hold | No | | `Provisioning` | Capacity acquired; the image, source, inputs and any saved state are being prepared | Yes | | `Running` | The command is running | Yes | | `Recovering` | The capacity was lost; the work is moving to new capacity | For new capacity, once acquired | | `Suspending` | State is being saved and compute released | Yes | | `Suspended` | Paused with its state saved; no compute is held | No compute | | `Cancelling` | Stopping and releasing compute | Until released | | `Succeeded` | Finished; outputs collected | No | | `Failed` | Finished without success; `status.reason` says why | No | | `Cancelled` | Stopped by `spec.state: Cancelled` or a delete | No | `Succeeded`, `Failed` and `Cancelled` are final: once there, the phase never changes and nothing more is billed. ## Transitions [Section titled “Transitions”](#transitions) | From | To | When | | ------------------------------------------------- | -------------- | ------------------------------------------------------------------------------------ | | (new) | `Queued` | The object is admitted | | `Queued` | `Provisioning` | Capacity is placed and acquired | | `Queued` | `Failed` | Nothing fits within `placement.queueTimeout` (`CapacityUnavailable`) | | `Provisioning` | `Running` | The command starts | | `Provisioning` | `Failed` | The container cannot start (`LaunchFailed`, `ImagePullFailed`) | | `Running` | `Recovering` | The capacity is lost | | `Recovering` | `Running` | The work restarts on new capacity | | `Recovering` | `Failed` | Recovery limits are used up (`RecoveryLimitExceeded`, `NoProgress`, `RestoreFailed`) | | `Running` | `Succeeded` | The work is done and its outputs are collected | | `Running` | `Failed` | The command failed more than `backoffLimit` allows, or `timeout` elapsed | | `Queued`, `Provisioning`, `Running`, `Recovering` | `Suspending` | `spec.state: Suspended`, credits or a budget ran out, or `maxCostUSD` was reached | | `Suspending` | `Suspended` | State is saved and compute released | | `Suspending` | `Running` | A suspend you asked for could not save state (`SuspendFailed`) | | `Suspending` | `Failed` | Saving state failed permanently or timed out | | `Suspended` | `Queued` | Resumed, and funds cover a new hold | | any unfinished phase | `Cancelling` | `spec.state: Cancelled`, or the object is deleted | | `Cancelling` | `Cancelled` | Compute is released and the final charge posted | ## spec.state [Section titled “spec.state”](#specstate) `spec.state` is what you want; `status.phase` is what is happening. `nodus suspend`, `nodus resume` and `nodus cancel` set it, and so can a manifest: | `spec.state` | Effect | | ------------------- | -------------------------------------------- | | `Running` (default) | Run, or resume from `Suspended` | | `Suspended` | Save state, release compute and stop billing | | `Cancelled` | Stop for good | A suspend caused by money (credits, a budget or `maxCostUSD`) leaves `spec.state` as it is and resumes on its own once the work is funded again or the cap is raised. Time spent suspended does not count against `timeout`. ## Conditions [Section titled “Conditions”](#conditions) | Condition | Meaning | | ------------------ | ------------------------------------------------------------------------------ | | `Admitted` | The object passed admission | | `Scheduled` | Capacity is placed; while false, its message says what it is waiting for | | `Funded` | Credits and budgets cover the work; false while it is stopped for money | | `Ready` | The command is running | | `Checkpointed` | The latest state save succeeded; false with `CheckpointFailed` when it did not | | `Suspended` | The work is suspended; `SuspendFailed` when a suspend could not save state | | `OutputsCommitted` | Every declared output is collected | | `SinksLoaded` | Every output sink has loaded into its table | Wait on a phase or a condition from the CLI: ```console $ nodus wait job/train --for=jsonpath='{.status.phase}'=Succeeded --timeout 2h $ nodus wait job/train --for=condition=OutputsCommitted ``` ## Pipelines and Sweeps [Section titled “Pipelines and Sweeps”](#pipelines-and-sweeps) A Pipeline or a Sweep takes its phase from its child Jobs. It is `Queued` until its first Job exists, `Running` while any runs, `Suspended` when every unfinished Job is suspended, and final once every Job is final: `Succeeded` if all succeeded, otherwise `Failed` with reason `StageFailed` (Pipeline) or `CellsFailed` (Sweep). Its `spec.state` passes to every unfinished child, and its `maxCostUSD` caps the spend of all its children together. # Container runtime contract > The directories, environment variables and sockets your code sees inside every Nodus container. Every Job, Sandbox and Function worker runs your image in an isolated container. The same contract holds on every offering: your code can rely on the paths and variables below wherever Nodus places it. ## Directories [Section titled “Directories”](#directories) | Path | Variable | What it is for | | --------------------------- | ---------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `/nodus/state` | `NODUS_STATE_DIR` | Recovery state. Write your model, optimizer and progress files here; Nodus checkpoints this directory and restores it on the next attempt. | | `/nodus/inputs/` | `NODUS_INPUT_` | Each declared input, read-only, downloaded before your command starts. The variable names the file (a URL or one object) or the directory (an object prefix, such as an earlier stage’s outputs). An input may set its own `path` instead. | | `/nodus/outputs` | `NODUS_OUTPUT_DIR` | Results. Every regular file here is uploaded when your command exits with code 0, with its SHA-256, and appears as the `outputs` output. Symbolic links are not collected. | | `/dev/shm` | | Shared memory for NCCL between GPUs and for data-loader workers: half the container’s memory limit, or half the machine’s memory without one. It counts against the memory limit. | | `/etc/nodus/hostfile` | | Multi-node Jobs (Beta): one `
slots=` line per node in rank order, for DeepSpeed and MPI tools. | | `/run/secrets//` | | Secret values as read-only files (mode `0400`) on a memory-backed filesystem, so they never reach a disk. The environment also carries each value, named by its key. | | `/run/nodus/events.sock` | `NODUS_EVENTS_SOCKET` | Progress, metrics and the checkpoint handshake (below). | | `/run/nodus/api.sock` | `NODUS_RUNTIME_SOCKET` | The Nodus API and model calls, authenticated as your Job, Sandbox or Function (below). | `NODUS_CHECKPOINT_DIR` is kept as another name for `NODUS_STATE_DIR` for existing programs. ## Environment variables [Section titled “Environment variables”](#environment-variables) | Variable | Set for | Meaning | | -------------------------------------------------- | -------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------- | | `NODUS_ATTEMPT` | Every container | The attempt id. A retried or recovered run gets a new attempt. | | `NODUS_JOB`, `NODUS_INDEX`, `JOB_COMPLETION_INDEX` | Jobs | The Job and, for indexed Jobs, this index. | | `NODUS_RESTORED` | Recovered attempts | `1` when `/nodus/state` was restored from a checkpoint before your command started. | | `NODUS_CURSOR_COMPLETED`, `NODUS_CURSOR_TOTAL` | Restartable Jobs | The progress cursor your program last reported. | | `NODUS_PARAM_` | Sweep cells | The cell’s parameters. | | `PET_NPROC_PER_NODE` | Single-node GPU Jobs | The GPU count, which `torchrun` reads as its `--nproc-per-node` default, so `torchrun train.py` starts one process per GPU. Your own value or flag wins. | Restoring a checkpoint restores files, never process memory: your program starts from the beginning and reads its own progress files from `/nodus/state`. Nodus never adds resume flags to your command. ## Network [Section titled “Network”](#network) Jobs, Functions and Workspaces reach the public Internet by default; Sandboxes and Agents have no outbound network unless their spec opens it. Either way a container never reaches private addresses (`10.0.0.0/8`, `172.16.0.0/12`, `192.168.0.0/16`), cloud metadata services, other containers or the machine it runs on, and has no IPv6. `localhost` always works inside the container. `/etc/resolv.conf` points at public resolvers, and `/etc/hosts` resolves `localhost`. With an allow list (`egress: AllowList`) the container has no direct route out. `HTTPS_PROXY` and `HTTP_PROXY` point at a proxy on the container’s loopback (`http://127.0.0.1:3128`) that reaches only the listed hosts, on ports 80 and 443. A `*.example.com` entry allows every subdomain of `example.com`. Most HTTP clients (`pip`, `npm`, `curl`, `git` over HTTPS, the Python and Node SDKs) use the proxy on their own. A listed name that resolves to a private address is still refused. CPU Sandboxes run under gVisor, whose `localhost` belongs to the sandbox alone, so there the proxy and the inference proxy listen on the sandbox’s gateway address instead of `127.0.0.1`. Read the address from `HTTPS_PROXY` and `OPENAI_BASE_URL` rather than writing `127.0.0.1` into your code; `NO_PROXY` already covers it. On a multi-node Job whose nodes share a private network, each rank’s `eth0` lists the node’s private address first, so NCCL, Gloo and `torchrun` advertise an address the other ranks reach. Ranks reach each other only on the rendezvous and NCCL ports, whatever the Job’s outbound setting. ## GPUs [Section titled “GPUs”](#gpus) A container on a GPU offering sees exactly the GPUs assigned to it, numbered from 0 as CUDA and PyTorch see them, with the NVIDIA driver libraries mounted read-only. It never sees another container’s GPUs. Use `nvidia-smi` or `torch.cuda.device_count()` to check what you have; do not set `CUDA_VISIBLE_DEVICES` yourself. ## Sidecars and the init command [Section titled “Sidecars and the init command”](#sidecars-and-the-init-command) Sidecars start before your command, in the order you list them, in the same container: they share its files, network, environment and logs. Each must answer its readiness probe (an HTTP `GET` of its path, any status below 400, or a TCP connection to its port) within 15 minutes before the next starts, and they stop when your command exits. A sidecar runs from your container’s image. `initCommand` runs after the sidecars and before your command, for preflight checks such as imports, free disk or the GPU count. Your command starts only if it exits with code 0; any other code fails the attempt with that code, and running past its timeout (30 minutes unless you set one) fails it with code 124. Both run inside your billed time. ## Commands you run in a container [Section titled “Commands you run in a container”](#commands-you-run-in-a-container) Commands you start in a running container (`nodus exec`, Processes) run with the same environment, secrets and directories as your main command, plus any variables you add. Their output is streamed to you and, for Processes, kept with your logs with secret values masked. A command started from a terminal session stops when you disconnect; a Process keeps running until it exits, you cancel it, its timeout passes or the container stops. Port forwarding and preview URLs reach a server listening on the container’s `localhost`, whatever its outbound network setting. In a CPU Sandbox they reach the sandbox’s own address instead, so listen on all interfaces (`0.0.0.0`); the same holds for a sidecar’s readiness port. File operations (`nodus cp`, the console Files tab) read, write and watch paths as your container sees them, including `/nodus/state` and `/nodus/outputs`, with the permissions of your container’s user: files written this way belong to that user and appear only once complete. `/proc`, `/sys` and `/dev` are not available to them. ## The events socket [Section titled “The events socket”](#the-events-socket) The socket speaks newline-delimited JSON, one object of at most 16 KiB per line. Send telemetry as objects with a `type` and, to make retries safe, a unique `id`: ```json {"id": "evt-10", "type": "progress", "completed": 120, "total": 1000} {"id": "evt-11", "type": "log.metrics", "step": 1200, "loss": 0.41} ``` When a program cannot reach the socket, it can print the same object on standard output after the prefix `nodus.event `. The line also stays in your logs. A program that sends no `log.metrics` events still gets its training metrics charted: Nodus reads loss, learning rate, epoch, step and `eval_*` values from Hugging Face Trainer and PyTorch Lightning log lines until the program sends a `log.metrics` event of its own. With `recovery.checkpoint.integration: HFTrainer`, the Trainer helper package is mounted at `/.nodus/python` and put first on `PYTHONPATH`, so its `sitecustomize` registers the Nodus callback in images without the Nodus SDK. ### Checkpoint handshake [Section titled “Checkpoint handshake”](#checkpoint-handshake) Nodus decides when to checkpoint. To save a consistent state first, send `{"type": "checkpoint.subscribe"}` once. Before each checkpoint Nodus writes a request on the connection: ```json {"type": "checkpoint.request", "requestId": "ck-1790000000000", "seq": 1790000000000, "urgent": true} ``` Finish writing your files to `/nodus/state`, then answer with `{"type": "checkpoint.ready", "requestId": "ck-1790000000000"}`. `urgent` means the capacity is about to go away: save quickly. A program that never subscribes is checkpointed without being asked, so write your state files atomically (write to a temporary name, then rename). A checkpoint of an empty state directory never replaces an earlier checkpoint that had files in it. ## The API socket [Section titled “The API socket”](#the-api-socket) `/run/nodus/api.sock` is HTTP over a Unix socket. Requests under `/apis/nodus.dev/` reach the Nodus API, and requests under `/v1/` (OpenAI- and Anthropic-compatible routes) reach Nodus inference. Both carry the service account token of the Job, Sandbox or Function the container belongs to: Nodus adds it to each request, so the token is never in your environment or files, and any `Authorization` header you send is replaced. ```sh curl --unix-socket "$NODUS_RUNTIME_SOCKET" http://nodus/apis/nodus.dev/v1/... ``` With the inference proxy enabled, the same model routes are also served on `http://127.0.0.1:7777` inside the container (on the gateway address in a CPU Sandbox), even with outbound network off, and `OPENAI_BASE_URL`, `ANTHROPIC_BASE_URL`, `OPENAI_API_KEY` and `ANTHROPIC_API_KEY` are set for it. The keys are placeholders, so SDKs and tools that take a base URL work unchanged. That port serves nothing but model calls. ## Signals and exit codes [Section titled “Signals and exit codes”](#signals-and-exit-codes) Your command runs under a small init process (`tini`, at `/.nodus/bin`) that forwards signals to your command’s process group and reaps finished child processes, so your command does not need to be written as PID 1. A stop sends `SIGTERM` to your command, then `SIGKILL` after the stop grace period (30 s unless your spec sets another). Exit code 0 completes the attempt; any other code, or a signal, fails it with your exit code and the last lines of your logs, with secret values masked. # Pricing > How Nodus sets your rate, what each second and token costs, and how to see the price before you launch. Nodus is prepaid: you buy credits, every run shows its rate before it starts, and the rate stays frozen for that run. The numbers on this page come from the same price list the API uses, so the [pricing reference](/docs/reference/pricing/) and your estimates always agree. This page explains how those numbers are set. ## See the price before you launch [Section titled “See the price before you launch”](#see-the-price-before-you-launch) `nodus run --dry-run` and the console review screen show the estimate for a run: the hourly rate, the expected cost range, the expected boot and teardown time as their own lines, and the minimum charge. The estimate is valid for up to 31 minutes. Launching with the estimate’s `ETag` holds you to it: if prices change in between, the launch is refused and you review the new estimate. ## Machines Nodus rents for you [Section titled “Machines Nodus rents for you”](#machines-nodus-rents-for-you) A Job, a GPU Sandbox, a GPU Workspace or a GPU Function worker runs on a machine Nodus rents for you. For that machine you pay exactly what the provider bills, divided by 0.875, so Nodus keeps 12.5 % of what you pay: * **From creation to deletion.** The clock runs from the moment the provider starts billing until the machine is confirmed deleted. Your usage is itemized by segment: `Boot` (start-up, image pull, readiness), `Restore` (resuming from a checkpoint), `Running` and `Teardown` (stop to deletion, plus the provider’s billing increment, charged once per machine). * **Never above list.** Every accelerator and count has a list price. Nodus never places your work on a machine whose rate would exceed it. * **Frozen for the run.** The rate is fixed when the machine is acquired. A later price change applies only to runs launched after it takes effect. Spare machines Nodus starts to finish your work sooner, and failures Nodus causes, are never charged to you. The list price is the ceiling; the “from” rate is the lowest rate a machine is available at right now, refreshed every minute. Your estimate shows the rate for your run. | Accelerator | ×1 list | ×1 from | ×8 list | ×8 from | | ------------- | ------- | --------- | ------- | --------- | | A10 | $0.89 | List only | $7.12 | List only | | A100 40G | $1.49 | List only | $11.92 | List only | | A100 40G PCIE | $1.39 | List only | $11.12 | List only | | A100 80G | $1.99 | List only | $15.92 | List only | | A100 80G PCIE | $1.89 | List only | $15.12 | List only | | B200 | $6.49 | List only | $51.92 | List only | | H100 PCIE | $2.99 | List only | $23.92 | List only | | H100 SXM | $3.29 | List only | $26.32 | List only | | H200 | $4.29 | List only | $34.32 | List only | | L4 | $0.89 | List only | $7.12 | List only | | L40S | $1.29 | List only | $10.32 | List only | | RTX 3090 | $0.59 | List only | $4.72 | List only | | RTX 4090 | $0.69 | List only | $5.52 | List only | | RTX 6000 ADA | $1.09 | List only | $8.72 | List only | | RTX A6000 | $0.79 | List only | $6.32 | List only | USD per hour for the whole machine. Pricebook 2026.10.8, effective 2026-10-01. ## Sandboxes, Functions and builds [Section titled “Sandboxes, Functions and builds”](#sandboxes-functions-and-builds) CPU Sandboxes, Functions, agent workers, image builds and CPU Jobs placed on Nodus nodes are billed per second from placement to release, at published rates per vCPU, per GiB of memory and per GiB of disk above 10 GiB per vCPU. The smallest billable shape is 0.25 vCPU with 512 MiB of memory; smaller requests are billed at that shape. ## Models [Section titled “Models”](#models) Each model’s price is the model’s cost divided by 0.95. When a model is served from more than one source, it is listed once, at the price of the cheapest source available now, and a request is charged the price it was accepted at even if another source finishes it. A request is priced once: every input, cached, output, audio or speech unit is added up exactly, then rounded up to the next micro-dollar. A model is served only for the operations it has a price for. If Nodus cannot confirm the outcome of a request, you are not charged for it. Indra (`nodus/indra`) picks one of 10 models for each request. On the Indra plan you pay the chosen model’s rates plus the routing call (0.044211 USD per 1M input tokens), in one price per request. Without the plan, free Indra routes among the three cheapest models at no charge, up to 100 requests and 200,000 tokens per organization per UTC day. The Indra plan costs 20.00 USD a month and includes 20.00 USD of `nodus/indra` usage each month at the listed prices. The allowance expires at the end of each month, and usage beyond it draws from your credits. Agents that run on Nodus-managed models are billed from your credits at the model’s cost divided by 0.875, on the agent run. The plan allowance does not apply to them. ## Storage, egress and credits [Section titled “Storage, egress and credits”](#storage-egress-and-credits) * **Storage:** 10 GB per org is included; beyond it, retained bytes are billed per GB-month. * **Egress:** 10 GiB per org per day is included; beyond it, egress is billed per GiB up to your daily egress quota. * **New organizations** receive 30.00 USD of credit, valid for 30 days. ## Rounding [Section titled “Rounding”](#rounding) Money is counted in micro-dollars. A running machine is charged as it goes, rounded down, and settled when it stops, rounded up once, so charging in many small windows costs the same as charging once. Current rates for every line are in the [pricing reference](/docs/reference/pricing/). # Sandbox isolation > What keeps a Sandbox apart from the machine, from other Sandboxes and from the network, and what is kept when it stops. A [Sandbox](/docs/guides/sandboxes/) runs code you did not write, such as the output of a model, so the boundary around it matters more than it does for your own jobs. This page lists the layers of that boundary, in the order a piece of code meets them. ## Where Sandboxes run [Section titled “Where Sandboxes run”](#where-sandboxes-run) Sandboxes run on **CPU machines Nodus operates**, not on GPU machines rented for your jobs. Nodus packs machines **per organization**: a machine serves one organization at a time, your Sandboxes share it with your other Sandboxes, agent workers and CPU jobs, and with nobody else’s. An empty machine is destroyed after 10 minutes and never handed to another organization, so one organization’s files and memory never sit on a machine another organization uses next. ## Layers [Section titled “Layers”](#layers) ### 1. A user-space kernel (gVisor) [Section titled “1. A user-space kernel (gVisor)”](#1-a-user-space-kernel-gvisor) Each Sandbox runs under [gVisor](https://gvisor.dev/) (`runsc`, in its `systrap` mode). gVisor answers the Sandbox’s system calls in its own kernel written in a memory-safe language and passes the host kernel only a small, filtered set of calls. A bug in the host kernel’s handling of an unusual system call is reachable from a Sandbox only through that narrow filter, not directly. ### 2. A network namespace and a firewall per Sandbox [Section titled “2. A network namespace and a firewall per Sandbox”](#2-a-network-namespace-and-a-firewall-per-sandbox) Every Sandbox gets its own network namespace with a default-deny `nftables` firewall: * Private ranges (RFC 1918), carrier-grade NAT, link-local addresses and cloud metadata addresses are always blocked, whatever the policy, so a Sandbox cannot reach the machine, its neighbours or the cloud’s metadata service. * Policy `Deny` (the default) has no route out at all. * Policy `Open` translates traffic to public addresses only. Nothing is allowed in from the network. Bytes in and out are counted, and a blocked connection raises an `EgressDenied` Event on the Sandbox. ### 3. Resource limits [Section titled “3. Resource limits”](#3-resource-limits) Each Sandbox runs in its own cgroup (v2) that limits CPU and memory to what you asked for and the number of processes to 256, so a fork bomb or a memory leak stops at the Sandbox’s own limits instead of slowing its neighbours. The root filesystem is backed by a per-Sandbox file under a disk quota (`resources.disk`), not by a directory shared with other Sandboxes. ### 4. An unprivileged process [Section titled “4. An unprivileged process”](#4-an-unprivileged-process) Commands run as a non-root user with no Linux capabilities and `no_new_privs` set, so the process cannot gain privileges through a set-uid program. The user comes from the image. The Nodus helper that serves commands and files runs outside the user’s process tree and is read-only to it. ## Continuity: what survives a stop [Section titled “Continuity: what survives a stop”](#continuity-what-survives-a-stop) `spec.continuity.mode` decides what a Sandbox keeps when it stops or when its machine is lost: | Mode | A stop keeps | If the machine is lost | | ----------------------- | ---------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------ | | `Snapshotted` (default) | The filesystem changes the Sandbox made, `/workspace` first among them | The Sandbox restarts on another machine from its last snapshot; work since that snapshot is gone | | `Ephemeral` | Nothing: the next start is empty | The Sandbox moves to `Failed` with the reason `NodeLost` | Processes never survive a stop: a start runs a fresh container from the image with the saved files restored, so start long-lived servers again from your code. Keep your work in `/workspace`, the directory the Sandbox starts in. `Snapshotted` is the right choice for an agent that works in a repository; `Ephemeral` suits a fresh sandbox per task. Snapshots are taken at every stop and about every 10 minutes while the Sandbox runs. ## Where gVisor is not available [Section titled “Where gVisor is not available”](#where-gvisor-is-not-available) Production Nodus CPU nodes ship with `runsc`, and Sandboxes there run under gVisor as described above. On a node without `runsc` (the local development provider on a laptop, or a Linux host where gVisor is not installed), `nodusd` runs the Sandbox under `runc` instead, with the same network namespace, firewall, cgroup limits and unprivileged process. Layers 2 to 4 are identical; layer 1 is not: a `runc` Sandbox shares the host’s Linux kernel, so it is a boundary for development and trusted code, not for hostile code. Caution Do not run untrusted code on a development node that lacks `runsc`. ## What a Sandbox does not protect against [Section titled “What a Sandbox does not protect against”](#what-a-sandbox-does-not-protect-against) * **Secrets you put in it.** Anything in a Sandbox’s environment or files is readable by the code running there. With `egress.policy: Open`, that code can send it anywhere on the internet. Keep `Deny` unless the work needs the network, and give a Sandbox only the credentials its task needs. * **Resource use up to your limits.** A Sandbox can use all the CPU and memory it asked for, and is billed for them while it runs. Set `maxCostUSD` to cap the spend of code you do not control. # Placement and scheduling profiles > How Nodus chooses where a run executes, what the estimate includes, and how profiles, deadlines and budgets change the choice. By default Nodus places each run on the **cheapest offering that can start it now**: the lowest hourly rate among the offerings that fit your request, have a machine free and stay within your limits. If that offering cannot be used, the next cheapest takes the run, and so on. You pay the rate of the offering the run starts on, shown in the estimate and frozen for the run. This page explains what the estimate includes and the settings that change the choice. ## Cost to completion [Section titled “Cost to completion”](#cost-to-completion) For every offering that fits your request, Nodus estimates the whole bill of the run. The estimate shows it for the chosen offering, and the `Cost` profile ranks offerings on it: * **Startup**: the machine’s boot, the image pull and any restore from a checkpoint. You pay for these because the capacity bills from the moment it is created. * **Running time**: from your declared `expectedDuration`, a training runtime’s measured estimate, or the history of earlier runs with the same labels on the same accelerator family. * **Teardown**: the time until the machine is confirmed deleted. * **Rounding**: each machine bills in increments, and the estimate rounds up the same way. * **Interruptions**: for interruptible capacity, the expected number of reclaims times the work redone after each one. Work that checkpoints loses only the stretch since its last save; work that saves nothing loses half its run on average. Under `Cost`, a cheap interruptible offering therefore wins only when its expected cost, lost work included, still beats on-demand capacity. A run whose checkpoints are all empty is costed as if it saves nothing. When the running time is unknown, every profile ranks offerings by hourly rate and the estimate shows the rate and the startup and teardown charge but no total. ## Profiles [Section titled “Profiles”](#profiles) `placement.profile` picks how the scheduler trades cost against time: | Profile | Chooses | | -------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `Balanced` (default) | The lowest hourly rate that meets your deadline and budget. Among offerings within 2 % of that rate, the healthier one, or the one that starts sooner, fits better or is already warm | | `Cost` | The lowest expected cost to completion, startup, teardown and lost work included, even if it is slower to start. Among offerings within 2 % of that cost, warm capacity and a better fit win | | `Speed` | The fastest expected finish among offerings within 1.5 × the cheapest expected cost | `Balanced` never chooses an offering whose rate is more than 2 % above the cheapest that meets your limits, and `Cost` never chooses one whose expected cost is more than 2 % above the cheapest at equal health; `interruptible: Prefer` gives interruptible capacity a 10 % allowance. Under `Cost` and `Speed`, an offering that often fails to start counts as dearer by that risk; under every profile, one that recently failed to start ranks lower for a few minutes within the 2 % band. When two offerings are within 2 % of each other, identical requests are spread across both instead of all taking the same one. ## Deadlines and budgets [Section titled “Deadlines and budgets”](#deadlines-and-budgets) * `placement.completeByTime`: offerings whose p90 finish is later are not used. * `maxCostUSD`, Budgets and your balance: offerings whose expected cost exceeds the money left are not used. * `placement.maxRateUSDPerHour`: offerings above this rate are not used. If nothing remains, the run waits in `Queued` and the estimate says why, for example `MissesDeadline` or `ExceedsRemainingBudget`. ## Estimates and the If-Match ceiling [Section titled “Estimates and the If-Match ceiling”](#estimates-and-the-if-match-ceiling) `--dry-run=server -o estimate` returns the expected cost p50 and p90, the startup time (cold, and warm when idle capacity of yours fits), the first hold, the minimum charge and `validUntil`, which is at most 31 minutes away. Creating with the estimate’s `If-Match` binds the launch to it: Nodus then uses no offering above the estimated rate plus 10 %, unless you set `placement.maxRateUSDPerHour` yourself. If prices moved beyond that, the run waits with `PriceAboveEstimate` instead of costing more than you saw. ## Reuse before new capacity [Section titled “Reuse before new capacity”](#reuse-before-new-capacity) Before buying new capacity, the scheduler considers capacity you already pay for: an idle machine of yours that fits reuses the increment already paid, and CPU work packs onto your Nodus nodes. A [BYOC pool](/docs/guides/pools/) named in `placement.pool` is used first at no hourly charge. ## Why a placement was made [Section titled “Why a placement was made”](#why-a-placement-was-made) `nodus describe` shows each attempt’s placement: the profile, the scores of the chosen offering (fit, time to result, cost to complete, recovery value, health), the fallbacks in order and every rejected offering with its reason. Offerings are shown by name, such as `h100-sxm-80g-x8-us`. ## Multi-node runs (Beta) [Section titled “Multi-node runs (Beta)”](#multi-node-runs-beta) For `spec.distributed`, Nodus first resolves the topology (nodes × GPUs per node) and then places every node together: * `network: Colocated` (the default) keeps every node in one location on one network; `Regional` keeps them in one region class; `Global` allows anywhere. * `transport: Direct` uses only private or direct paths between nodes; `Auto` also allows the relayed mesh, with lower bandwidth. * The whole gang is priced together, from the slowest node’s start. The estimate also shows the assembly bound: the most a failed assembly can cost. If no set of offerings satisfies these rules, the run waits with `GangInfeasible`. See [multi-node training](/docs/guides/multi-node/). # Install the CLI > Install the nodus CLI on macOS, Linux or Windows, verify the download, turn on shell completion and sign in. The `nodus` CLI is one self-contained binary for macOS, Linux and Windows on x86-64 and ARM64. Pick one way to install it. * pip ```sh pip install nodus-compute ``` The Python package (Python 3.10 or newer) includes the CLI, so `nodus` is on your `PATH` wherever the package is installed. Use this if you also want the Python SDK. * Homebrew ```sh brew install --cask nodus-compute/tap/nodus ``` * macOS and Linux ```sh curl -fsSL https://nodus-compute.ai/install | sh ``` The script installs to `~/.local/bin` and checks the archive against the release’s `checksums.txt` first. `NODUS_VERSION=1.2.3` pins a release and `NODUS_INSTALL_DIR` picks another directory. * Windows ```powershell irm https://nodus-compute.ai/install.ps1 | iex ``` Check that it works: ```console $ nodus version nodus v1.0.0 ``` ## Sign in [Section titled “Sign in”](#sign-in) ```sh nodus login ``` Your browser opens to confirm the sign-in. If you belong to several orgs, pick the ones this machine should use: the CLI stores one API key per org in the OS keychain and creates one **context** per org. Switch between them with `nodus config use-context `, or pass `--org ` to a single command. | Where you are | Command | | --------------------------------------------- | -------------------------------------------------------------------------------- | | A laptop with a browser | `nodus login` | | An SSH session or a machine without a browser | `nodus login --device`, then approve the code from any device | | CI, with a key in a secret | `echo "$NODUS_API_KEY" \| nodus login --with-token`, or just set `NODUS_API_KEY` | `nodus whoami` shows who you are signed in as, your role, the current project and your available credit. `nodus logout` revokes this machine’s key and removes the context. Environment variables The CLI reads six variables, all optional: `NODUS_API_KEY`, `NODUS_API_URL`, `NODUS_ORG`, `NODUS_PROJECT`, `NODUS_CONTEXT` and `NODUS_CONFIG`. A variable wins over the config file for that invocation. ## Shell completion [Section titled “Shell completion”](#shell-completion) ```sh nodus completion zsh > "${fpath[1]}/_nodus" # zsh nodus completion bash > /etc/bash_completion.d/nodus # bash (or ~/.local/share/bash-completion/completions/nodus) nodus completion fish > ~/.config/fish/completions/nodus.fish ``` Completion covers every command and flag. ## Verify a download [Section titled “Verify a download”](#verify-a-download) Every release publishes `checksums.txt`, a keyless cosign signature over it, and an SBOM per archive. To verify an archive you downloaded yourself: ```sh cosign verify-blob checksums.txt \ --signature checksums.txt.sig --certificate checksums.txt.pem \ --certificate-identity-regexp '^https://github.com/nodus-compute/nodus-platform/' \ --certificate-oidc-issuer https://token.actions.githubusercontent.com shasum -a 256 --ignore-missing -c checksums.txt ``` ## Where the CLI keeps things [Section titled “Where the CLI keeps things”](#where-the-cli-keeps-things) | Path | What it holds | | -------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------- | | `~/.nodus/config` | Contexts: API server, org and project. It is a kubeconfig, so `KUBECONFIG=~/.nodus/config kubectl get jobs.nodus.dev` works too | | OS keychain (`~/.nodus/credentials`, mode 0600, where there is none) | One API key per context | | `~/.nodus/cache/` | The cached list of resource kinds; `nodus api-resources` refreshes it | If you used Nodus before 1.0, its `~/.nodus/config.toml` is renamed to `config.0x.bak` the first time the new CLI runs; sign in again with `nodus login`. A 0.x `nodus.toml` converts into a manifest you can review and apply: ```sh nodus convert nodus.toml > job.yaml # fields that do not carry over are listed on stderr nodus apply -f job.yaml --dry-run=server -o estimate ``` ## Upgrade and uninstall [Section titled “Upgrade and uninstall”](#upgrade-and-uninstall) Upgrade the same way you installed (`pip install -U nodus-compute`, `brew upgrade --cask nodus`, or run the install script again). To uninstall, run `nodus logout`, remove the binary, and delete `~/.nodus`. ## Next steps [Section titled “Next steps”](#next-steps) * [Quickstart](/docs/getting-started/quickstart/): run your own code on a GPU and download its results. * [Nodus for kubectl users](/docs/getting-started/kubectl-users/): the verbs you already know. * [CLI reference](/docs/reference/cli/): every command and flag. # Nodus for kubectl users > How Nodus maps onto the Kubernetes API model, which kubectl verbs nodus supports, and how to point kubectl itself at Nodus. If you know kubectl, you already know most of `nodus`. Every Nodus resource (Jobs, Sandboxes, Volumes, Secrets, InferenceEndpoints, Budgets and the rest) is a Kubernetes-style object in the `nodus.dev` API group, with `metadata`, `spec` and `status`, and the CLI speaks the same verbs over all of them. ## The mapping [Section titled “The mapping”](#the-mapping) | Kubernetes | Nodus | | -------------------- | ------------------------------------------------------------------------------------ | | Cluster | The Nodus API (`https://api.nodus-compute.ai`) | | Namespace (`-n`) | Project (`-p`; `-n` is accepted) | | User and credentials | An API key per org, kept in the OS keychain | | kubeconfig context | One context per org, in `~/.nodus/config` | | `kubectl get pods` | `nodus get jobs`, `nodus get sandboxes`, `nodus get all` | | `kubectl exec` | `nodus exec` into a Job, Sandbox or Workspace; each command is recorded as a Process | Your org comes from the API key, so there is no org in a manifest. Projects are created in the console or with `nodus create -f`, and `default` always exists. ## Verbs you already know [Section titled “Verbs you already know”](#verbs-you-already-know) ```sh nodus get jobs -o wide # tables are rendered by the server nodus get jobs -l team=nlp --field-selector status.phase=Running nodus get job/train -o jsonpath='{.status.phase}' nodus get jobs -o custom-columns=NAME:.metadata.name,GPU:.spec.resources.gpu nodus get jobs -w # watch nodus describe job/train # includes a Placement section and events nodus apply -f job.yaml # client-side three-way merge nodus diff -f job.yaml # what apply would change (exit 1 on differences) nodus apply -f jobs/ --prune -l app=nightly # delete what was applied before and is gone now nodus edit job/train # $EDITOR, saved as a merge patch nodus patch job/train --type merge --patch '{"spec":{"maxCostUSD":"50"}}' nodus label job/train tier=gold nodus wait job/train --for=jsonpath='{.status.phase}'=Succeeded --timeout 30m nodus delete job/train --wait nodus logs -f job/train nodus exec -it sb/dev -- bash nodus shell sandbox/dev # create the Sandbox if needed, then exec -it bash nodus events -w --for job/train # events as they are recorded nodus port-forward job/train 6006 nodus explain job.spec.resources nodus api-resources nodus auth can-i create jobs ``` Short names work as in kubectl: `sb` for sandboxes, `vol` for volumes, `sec` for secrets, `sa` for service accounts. `nodus api-resources` lists them all. ## What is different [Section titled “What is different”](#what-is-different) * **Server dry-run returns a cost estimate.** `nodus apply -f job.yaml --dry-run=server -o estimate` prints the estimated cost, start time and the hold that would be placed, without creating anything. * **State verbs instead of scaling.** `suspend`, `resume` and `cancel` apply to Jobs, Pipelines and Sweeps; `start` and `stop` to Sandboxes, Workspaces, Functions, Agents and InferenceEndpoints. They set `spec.state`. * **Requests are annotations.** `nodus request restart sb/dev` (or `nodus rollout restart sb/dev`) asks a controller to act once, recorded on the object. * **Typed generators.** `nodus create sandbox dev --cpu 2`, `nodus create secret hf --from-literal HF_TOKEN=…`, `nodus create volume weights --size 200Gi`, `nodus create budget research --limit 2000`. Add `--dry-run=client -o yaml` to print the manifest instead of creating it. Kinds that show a secret once print it alone on standard output: `nodus create apikey ci --scopes jobs:write`, `nodus create enrollmenttoken --pool lab` (the host installer) and `nodus create token sa/ci`. `nodus create sshkey --from-file ~/.ssh/id_ed25519.pub` adds a public key, or `--generate` makes the pair. * **Outputs are files you download.** `nodus cp job/train:outputs/model ./model` copies a declared output and checks its SHA-256 digest. * **Merge patches only.** Strategic merge patch and server-side apply are not supported; `apply` does the three-way merge in the client and records the applied manifest in a last-applied annotation on the object. * **Exit codes.** Every command exits `0`, `1` on an error or `2` on a usage error, like kubectl. `nodus run` passes your command’s exit code through instead. ## Use kubectl itself [Section titled “Use kubectl itself”](#use-kubectl-itself) `~/.nodus/config` is a real kubeconfig whose user entry runs `nodus auth token` as an exec credential plugin, so kubectl and any client-go tool (k9s included) can read and write Nodus resources: ```sh export KUBECONFIG=~/.nodus/config kubectl get jobs.nodus.dev kubectl apply -f job.yaml kubectl get jobs.nodus.dev -w kubectl wait jobs.nodus.dev/train --for=jsonpath='{.status.phase}'=Succeeded ``` Use the full `jobs.nodus.dev` resource name with kubectl, because it also knows the built-in `batch/v1` Jobs. What kubectl cannot do against Nodus kubectl’s `logs`, `exec` and `port-forward` only work on Pods, so use `nodus` for those. There are no Pods, Namespaces or other core Kubernetes resources, and `kubectl auth can-i` is not served: use `nodus auth can-i` instead. ## Plugins [Section titled “Plugins”](#plugins) Like kubectl, `nodus` runs any executable called `nodus-` on your `PATH` as `nodus `, so you can add your own commands. # Nodus for Modal users > Move a Modal app to Nodus. The Python SDK keeps Modal's names wherever the concept matches, so most code changes by one import. The Nodus Python SDK keeps Modal’s names wherever the concept is the same: `App`, `@app.function`, `.remote()`, `.map()`, `.spawn()`, `@app.cls` with `enter` and `exit` hooks, `Image`, `Volume`, `Secret` and `Sandbox`. Most apps move by changing `import modal` to `import nodus`. This page lists what is the same, what is spelled differently and what Nodus adds. ## Sign in and run [Section titled “Sign in and run”](#sign-in-and-run) ```console pip install nodus-compute nodus login # opens the console, stores a key for each org you pick nodus run app.py # like `modal run`: an ephemeral App, deleted when the entrypoint returns nodus deploy app.py # like `modal deploy`: a persistent App nodus serve app.py # like `modal serve`: redeploys when a file changes ``` In CI, set `NODUS_API_KEY` instead of running `nodus login`. examples/python/quickstart/app.py ```python """Quickstart: one Function called three ways. Run it with `nodus run examples/python/quickstart/app.py --n 10`. """ import nodus app = nodus.App("quickstart") @app.function(cpu=1, memory="1Gi", max_cost=1) def square(x: int) -> int: return x * x @app.local_entrypoint() def main(n: int = 10) -> None: print("remote:", square.remote(7)) # one call; blocks for the result call = square.spawn(8) # start without waiting print("spawned:", call.get(timeout=600)) print("map:", list(square.map(range(n)))) # one call per input, results in input order ``` ## Side by side [Section titled “Side by side”](#side-by-side) | Modal | Nodus | | ------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------- | | `import modal` | `import nodus` | | `app = modal.App("x")` | `app = nodus.App("x")` | | `@app.function(gpu="H100", timeout=3600)` | `@app.function(gpu="H100", timeout="1h")` (numbers are also seconds) | | `f.remote(x)`, `f.map(xs)`, `f.spawn(x)`, `call.get()` | The same | | `f.local(x)` | The same | | `@app.local_entrypoint()` | The same | | `modal.Function.from_name("app", "f")` | `nodus.Function.from_name("app", "f")` (`Function.lookup` also works) | | `@app.cls()` with `@modal.enter()`, `@modal.method()`, `@modal.exit()` | `@app.cls()` with `@nodus.enter()`, `@nodus.method()`, `@nodus.exit()` | | `modal.Image.debian_slim().pip_install("torch")` | `nodus.Image.debian_slim().pip_install("torch")` | | `modal.Image.from_registry(...)`, `.apt_install`, `.run_commands`, `.env`, `.add_local_dir` | The same | | `modal.Volume.from_name("v", create_if_missing=True)` | The same; `vol.commit()` and `vol.reload()` behave as in Modal | | `modal.Secret.from_name("hf")`, `Secret.from_dict({...})` | The same, plus `Secret.from_dotenv(".env")` | | `modal.Sandbox.create(app=app, image=...)` | `nodus.Sandbox.create(image=...)` (no App needed) | | `sb.exec("python", "-c", "...")`, `p.stdout.read()`, `p.wait()` | The same | | `sb.open(path, "w")` | The same | | `sb.tunnels()` | Returns a list of `Tunnel(port, url, public)`; `sb.tunnels.open(8080)` is not available yet | | `sb.snapshot_filesystem()` | The same (Beta) | | `@modal.experimental.clustered(size=2)` | `@nodus.clustered(size=2)` (Beta) | | `min_containers`, `max_containers`, `scaledown_window` | The same, or `min_workers` and `max_workers` | | `await f.remote.aio(x)` | The same: every blocking call has an `.aio` form | GPU strings use Modal’s spellings: `"H100"`, `"H100:2"`, `"A100-80GB"`, `"A10G"`, `"L40S"`. A family such as `H100` matches any of its variants; `"H100!"` pins the exact variant. A list such as `["H100", "H200"]` accepts either. ## What Nodus adds [Section titled “What Nodus adds”](#what-nodus-adds) Nodus places every call on the cheapest capacity that finishes it on time, and it stops work before money runs out. These arguments have no Modal equivalent: ```python @app.function( gpu="H100", max_cost=40, # a hard cap in USD across this Function's workers checkpoint="/nodus/state", # files here are saved and restored if capacity is reclaimed interruptible=True, # allow cheaper interruptible capacity; progress is kept through the checkpoint region=["us", "eu"], # region classes, not provider regions ) def train(lr: float) -> dict: ... print(train.estimate(3e-4)) # dry-run: expected cost, cold and warm start, the hold it needs ``` * `f.estimate(...)` returns the expected cost and start time of one call before you run it. * A cold start on a GPU with no warm worker prints its expected wait, so a long first call is not a surprise. * Errors are typed: `nodus.errors.InsufficientCredits` states the amount needed and how to add credit. * `nodus.Job`, `nodus.Workspace` and `nodus.llm` cover batch jobs, development machines and inference with the same credentials. ## Differences to know [Section titled “Differences to know”](#differences-to-know) * **Parametrized classes** (`modal.parameter()`) are not supported. Configure the class in its `@nodus.enter()` hook instead. Calling `MyClass(arg=...)` raises `nodus.errors.Unsupported`. * **The Python minor version** of the image must equal yours, as in Modal. `Image.debian_slim()` defaults to your version; a mismatch is refused before anything runs. * **Web endpoints** (`@modal.web_endpoint`, `@modal.asgi_app`) are not available. Use `nodus.InferenceEndpoint` for model serving. * **`modal.Dict` and `modal.Queue`** have no equivalent. Pass data through return values, a Volume or your own database. * **Clustered Functions** (Beta) run each `.remote()` or `.spawn()` as one gang; `.map()` over a clustered Function raises `nodus.errors.Unsupported`. * **Timeouts and durations** accept Go-style strings (`"90s"`, `"6h"`) as well as seconds. Days are not a unit. * **Money** is always a decimal amount in USD (`max_cost=40` or `"40.00"`). ## Next steps [Section titled “Next steps”](#next-steps) * [Python SDK guide](/docs/guides/python/) * [Functions and classes](/docs/guides/python/functions/) * [Sandboxes](/docs/guides/python/sandboxes/) # Run your own code > Run a training script from your own directory on a GPU with nodus run, follow it, download its output and read what it cost. This page runs a script from your own directory on a GPU, then downloads the file it wrote. It assumes you have [installed the CLI and signed in](/docs/getting-started/install/). Every command here runs in CI as the `cli/quickstart` example. 1. **Put your code in a directory.** Any directory works; this one holds a small PyTorch training loop that saves its final metrics to `/nodus/outputs/metrics.json`. train.py ```python """Fit a small linear model, print progress and save metrics.json as the declared output "metrics".""" import json import pathlib import torch device = "cuda" if torch.cuda.is_available() else "cpu" print("device:", torch.cuda.get_device_name() if device == "cuda" else "cpu") torch.manual_seed(0) x = torch.randn(1024, 16, device=device) y = x @ torch.randn(16, 1, device=device) w = torch.zeros(16, 1, device=device, requires_grad=True) opt = torch.optim.SGD([w], lr=0.1) for step in range(200): loss = ((x @ w - y) ** 2).mean() opt.zero_grad() loss.backward() opt.step() if step % 50 == 0: print(f"step {step} loss {loss.item():.4f}") print(f"final loss {loss.item():.6f}") out = pathlib.Path("/nodus/outputs") out.mkdir(parents=True, exist_ok=True) (out / "metrics.json").write_text(json.dumps({"final_loss": loss.item(), "steps": 200})) ``` 2. **Run it.** From that directory: ```sh nodus run --name cli-quickstart --gpu L4 --image nodus/pytorch \ --output metrics=/nodus/outputs/metrics.json -- python train.py ``` ```console job/cli-quickstart created · est. $0.01–0.03 · starts in ~2–4 min (cold) · hold $0.50 · Ctrl+C to cancel, -d to detach ✓ Scheduled l4-24g-x1-us · $0.52/h (rate frozen) ✓ Provisioning 1m48s ▶ Running device: NVIDIA L4 step 0 loss 15.8732 ... final loss 0.000001 ✓ Succeeded in 2m31s · $0.02 (boot $0.01 · running $0.01) · kept warm 60 s · nodus describe job/cli-quickstart ``` Before anything is charged you see the estimate and a hold: the most this run can cost before you are asked. `nodus run` then uploads the directory (files listed in `.gitignore` or `.nodusignore` stay local), prints each phase, streams your program’s output on stdout and ends with the cost line. The CLI exits with your command’s exit code, so it works in scripts and CI. 3. **Get the output.** The `--output` flag declared `metrics`; copy it back: ```sh nodus cp job/cli-quickstart:outputs/metrics metrics.json ``` The file is checked against its SHA-256 digest and only written when it matches. ## Leave it running [Section titled “Leave it running”](#leave-it-running) Press **Ctrl+C** during a run and the CLI asks whether to cancel the Job; answer `n` to detach and leave it running. Or start with `-d` to return as soon as the Job exists. A detached Job keeps running: ```sh nodus get jobs # every Job in the project, with phase and cost so far nodus get jobs --mine -w # only yours, updating live nodus logs -f job/cli-quickstart # stream the output again nodus describe job/cli-quickstart # attempts, placement, cost, conditions and events nodus cancel job/cli-quickstart # stop it; you pay only for what ran ``` ## Control the cost [Section titled “Control the cost”](#control-the-cost) | Flag | Effect | | -------------- | ------------------------------------------------------ | | `--max-cost 5` | The Job stops gracefully before it spends more than $5 | | `--timeout 2h` | The Job is stopped after two hours of wall-clock time | | `--dry-run` | Print the estimate and exit without creating anything | `nodus billing` shows your balance and this month’s spend; `nodus billing usage --group-by day` breaks it down. Exit codes `nodus run` exits with your command’s own exit code. It exits `125` when Nodus could not run the command (an API error or an invalid request), `124` on `--timeout` and `130` when you interrupt it. ## Clean up [Section titled “Clean up”](#clean-up) Finished Jobs are kept for 30 days so you can read their logs and outputs, then deleted. To delete one now: ```sh nodus delete job/cli-quickstart ``` ## Next steps [Section titled “Next steps”](#next-steps) * [Nodus for kubectl users](/docs/getting-started/kubectl-users/): manifests, `apply`, `get -o`, `wait` and `diff`. * [CLI reference](/docs/reference/cli/): every command and flag, with tested examples. # API keys and scopes > Create API keys for scripts and CI, limit what they can do with scopes and projects, and revoke them. An API key lets a script, a notebook or CI call Nodus as you. Keys look like `nodus_sk_live_…` and are sent as `Authorization: Bearer `, or through `NODUS_API_KEY`. ```bash nodus create apikey notebook ``` The key is printed **once**, alone on standard output, so `KEY=$(nodus create apikey notebook)` captures it. Nodus stores only a keyed hash of it, so a lost key cannot be shown again: delete it and create a new one. `-o json` or `-o yaml` prints the whole object, key included. ## Scopes [Section titled “Scopes”](#scopes) A key’s scopes limit what it can do. The default, `*`, means “whatever my role allows”, and it follows your role: if your role changes, the key changes with it. ```bash nodus create apikey ci --scopes jobs:write,volumes:read --projects research --expires 720h nodus auth can-i create sandboxes # run with the key to check it ``` Scopes are `:read` or `:write` (write includes read), plus `*:read` for read-only access to everything. A key can never hold more than its creator’s role: asking for more is refused with `403` naming the scopes you lack. A key created with another key (or by a connected agent) can hold only what that credential holds, so it cannot ask for `*` or `*:read` unless the creating credential has them itself. `--scopes` and `--projects` take comma-separated lists, and `--description` says what the key is for. To mint a key for CI that survives its creator, bind it to a service account of the project you work in with `--service-account NAME`. A key restricted with `--projects` reaches only those projects, and it cannot change org-level settings. If every project it was restricted to is deleted, it reaches no project at all. A key bound to a service account can only be restricted to that service account’s project. ## List and revoke [Section titled “List and revoke”](#list-and-revoke) ```bash nodus get apikeys nodus delete apikey ci ``` A deleted key stops working within 30 seconds everywhere. Listing shows the key’s prefix (`nodus_sk_live_01j9…`), scopes, owner, last use and expiry, never the key itself. Members can delete their own keys. Deleting another member’s key takes an Admin or Owner, and deleting a service account’s key takes permission to manage service accounts in its project. ## Keys the CLI creates [Section titled “Keys the CLI creates”](#keys-the-cli-creates) `nodus login` creates one key per org, named `cli--`, with every scope your role grants, valid for 90 days and labelled as launched by the CLI. `nodus logout` revokes it. ## If a key leaks [Section titled “If a key leaks”](#if-a-key-leaks) Delete it. Keys are registered with GitHub secret scanning, so a key pushed to a public repository is reported to us and revoked. # Connected agents > See and disconnect the MCP clients you let act on Nodus for you. When an MCP client such as Claude Code, Cursor or Codex connects to Nodus, the consent page asks you to pick an org and the scopes the client may use. Nodus records that choice as an **OAuth grant**. The client can do only what the grant allows, and never more than your own role in that org. It cannot create API keys or other grants. A client is connected to one org at a time. Approving it again, for any org, replaces the earlier grant. ## Seeing what is connected [Section titled “Seeing what is connected”](#seeing-what-is-connected) Console › Settings › Connected agents lists each client with its scopes and when it was last used. Or: ```bash nodus get oauthgrants ``` You see your own grants. Admins and Owners see every member’s grants in the org. ## Disconnecting [Section titled “Disconnecting”](#disconnecting) ```bash nodus delete oauthgrant claude-code-3fa2c1 ``` The client stops working within 30 seconds, even with an access token it already holds. A grant also lapses 30 days after its last use, and removing a member from the org disconnects all of their clients there. # Device login > Sign the CLI in on a machine without a browser, or from CI with an existing key. On a remote server or container with no browser: ```bash nodus login --device ``` The CLI prints a code like `BDWP-HMTR` and a link. Open the link on any device where you are signed in, check that the code matches, pick the orgs and approve. The CLI finishes on its own within a few seconds. The code works for 10 minutes; denying it stops the login. The CLI polls every 5 seconds. Every key it receives is the same kind `nodus login` creates: one per chosen org, 90 days, revocable with `nodus logout` or `nodus delete apikey`. ## In CI [Section titled “In CI”](#in-ci) Pass an existing key instead of signing in: ```bash echo "$NODUS_API_KEY" | nodus login --with-token ``` Prefer a [service account token](/docs/guides/access/service-accounts/) for CI, so it keeps working when people leave. # Members and roles > Invite people to your org, choose their role, and remove members safely. ## Roles [Section titled “Roles”](#roles) | Role | Can | | ---------- | ------------------------------------------------------------------------------------------------- | | **Owner** | Everything, including granting or removing the Owner role | | **Admin** | Everything except the Owner role: members, invites, projects, API keys, service accounts, billing | | **Member** | Create and manage work: jobs, sandboxes, volumes, secrets, their own API keys | | **Viewer** | Read everything, change nothing | `nodus auth can-i create jobs` asks the server whether you hold a permission. ## Invite someone [Section titled “Invite someone”](#invite-someone) ```bash nodus create invite --email ada@example.com --role Member ``` The invite link works for 72 hours and only once. The invitee signs in with that email address (it must be verified) and accepts it in the console, where pending invites also appear as a banner. You can invite up to your seat limit, 10 members by default; pending invites count as seats. An invite nobody accepts stays under **Team › Invites** after it expires, marked Expired and holding no seat, so you can resend it. Once your org has bought credits, an Owner or Admin can change the seat limit under **Team › Members › Change**, up to 1,000. When an invite is accepted, or someone is made an Admin or Owner, every Owner and Admin gets an email. ```bash nodus get invites nodus request resend invite/inv-01j9abc # a fresh link, at most 3 times, 10 minutes apart nodus delete invite inv-01j9abc ``` You can invite someone with a role up to your own: Admins invite Admins, Members and Viewers; only Owners invite Owners. ## Let your company join without invites [Section titled “Let your company join without invites”](#let-your-company-join-without-invites) An Owner or Admin adds your company’s email domain under **Team › Company domains**, then adds the TXT record shown there at your DNS provider and selects **Verify**. After that, anyone who signs up with a verified email at that domain is offered your org during sign-up and joins it as a Member in one click. Seats still apply, public mailbox domains such as `gmail.com` are refused, and a domain belongs to one org at a time. Remove the domain to stop new joins; people who already joined stay members. Someone you remove from the org can only come back through an invite. ## Change a role [Section titled “Change a role”](#change-a-role) ```bash nodus get members nodus edit member usr-01j9abc --role Admin ``` An org always keeps at least one Owner: demoting or removing the last Owner is refused with `409 Conflict`. ## Remove a member [Section titled “Remove a member”](#remove-a-member) ```bash nodus delete member usr-01j9abc nodus delete member usr- # leave an org yourself ``` Removing a member revokes their API keys and third-party app access in that org at once. Keys bound to a [service account](/docs/guides/access/service-accounts/) keep working, so CI does not break when the person who set it up leaves. # Account security > Sign in and control access with organization roles and scoped credentials. Sign in with your email and password or an enabled Google or GitHub account. You can then use the resources and settings your organization role allows. Nodus does not require QR scanning, an authenticator app or a separate verification code after sign-in. Email verification for new accounts and password recovery still apply. CLI and device login use your existing signed-in session. API keys and service account tokens keep their own scopes. Review them under [API keys](/docs/guides/access/api-keys-and-scopes/). # Projects > Group work inside an org with projects, and restrict members and keys to some of them. A **project** is a namespace inside your org. Jobs, Sandboxes, Volumes, Secrets and the other resources you create live in one project. Every org has `default`, which cannot be deleted. ```bash nodus create project research --display-name "Research" nodus get projects nodus run -p research -- python train.py ``` In the console, **Team › Projects** lists each project with how many members can use it, its spend over the last 30 days and its spend limit, and has **New project**. Give each team its own project: its runs, storage and keys stay apart, and its spend shows on its own line. A project name is lowercase letters, digits and dashes. `nodus`, `team`, `settings`, `billing`, `integrations`, `infrastructure` and `support` are reserved. Creating, changing and deleting projects needs the `projects:write` scope, which Admins and Owners hold. Deleting a project deletes everything in it: running work is cancelled first. Usage and billing records are kept. ## Restrict access to projects [Section titled “Restrict access to projects”](#restrict-access-to-projects) Members and API keys can be limited to some projects: ```bash nodus edit member usr-01j9abc --projects research nodus create apikey ci --projects research --scopes jobs:write ``` A member or key with a project list sees and changes only those projects. An empty list means every project. ## Cap a team’s spend [Section titled “Cap a team’s spend”](#cap-a-teams-spend) A budget scoped to a project caps that team. From **Team › Projects**, open a project’s menu and choose **Set a spend limit**, or: ```bash nodus create budget vision-monthly --limit 500 --scope-project vision-team ``` Budgets can also cap one person, or one person inside one project. When several budgets cover the same work, the tightest one applies. # Quotas > See the limits on your org, what counts against them, and how to raise them. Quotas keep one mistake from running away with your credits. See yours with: ```bash nodus get quota default -o yaml ``` | Quota | Default before your first purchase | After a purchase | | --------------------------------------------------------------------- | ---------------------------------- | ---------------- | | `members` (seats, including pending invites) | 10 | 10 | | `apiKeys` (live keys, including CLI logins and ServiceAccount tokens) | 50 | 200 | | `liveSandboxes` (stopped ones count until their delete completes) | 10 | 100 | | `storageGiB` | 50 | 1024 | | `egressGiBPerDay` | 20 | 100 | | `nodusNodes` | 1 | 10 | | `gangNodes` (multi-node jobs) | 0 | 8 | | `inferenceRPM` / `inferenceTPM` | 60 / 100,000 | 600 / 1,000,000 | | `assistantUSDPerDay` | $1 | $10 | Going over a count quota refuses the create with `429 QuotaExceeded`, naming the quota. Nothing already running is touched. Daily quotas (`egressGiBPerDay`, `assistantUSDPerDay`) reset at 00:00 UTC. Expired and revoked API keys do not count against `apiKeys`. Multi-node jobs are in Beta: your org needs the distributed training beta, and `gangNodes` counts the nodes of running multi-node jobs. A job that would go over it waits in the queue instead of failing. Owners and Admins of an org that has bought credits set their own seats, up to 1,000, under **Team › Members › Change**, or with `PATCH /apis/nodus.dev/v1/quotas/default` and `{"hard": {"members": 25}}`. Seats never go below the members and pending invites you have. To raise any other org limit, or seats past 1,000, contact support. To cap what one project or one person spends, use a [Budget](/docs/guides/billing/#cap-spending-with-a-budget). # Service accounts > Give CI and code running on Nodus an identity that belongs to the project, not to a person. A **service account** is an identity for machines. Every project has one named `default`, used by code running inside Nodus for its own calls. Create more for CI and automation: ```bash nodus create serviceaccount ci -p research --scopes jobs:write,volumes:read ``` Managing service accounts needs `serviceaccounts:write` (Admins and Owners). ## Tokens for CI [Section titled “Tokens for CI”](#tokens-for-ci) ```bash nodus create token sa/ci -p research --duration 720h ``` This prints a key bound to the service account, once. It is valid for 90 days by default and at most one year. Its permissions are the service account’s scopes, narrowed further by any `--scope` you pass. Service account keys do not belong to a person: they keep working when the member who created them leaves the org, and they stop working when the service account is deleted. Use them for the GitHub Action and any shared automation. ## Inside Nodus [Section titled “Inside Nodus”](#inside-nodus) Code running in a Job, Sandbox or Agent reaches the API through `/run/nodus/api.sock` with a short-lived token of its service account. The token is never in the environment or on disk, and it stops working the moment its run is stopped or moved. # Sign up and create your org > Create an account, create or join an organization, and switch between the orgs you belong to. Everything you run on Nodus belongs to an **organization** (org). Credits, projects, members and API keys are all per org. ## Sign up [Section titled “Sign up”](#sign-up) Sign up at the console with email, Google or GitHub. Nothing is created until you choose: * If someone invited you, the console shows the invite first. Accept it to join their org. * Otherwise the console asks you to **Create your org**. The name is filled in for you, and you can change it. The first org you create with a verified email address receives the **$30 starter credit**, valid for 30 days. Joining an org by invite never adds credit, so accepting an invite does not use up your own starter credit. ## Start from the terminal [Section titled “Start from the terminal”](#start-from-the-terminal) ```bash nodus login ``` This opens the console in your browser. Sign in, pick the orgs this computer should use, and the CLI stores one key per org in your OS keychain. If you have no org and no pending invite, `nodus login` creates your org for you, with the starter credit. On a machine without a browser, use [device login](/docs/guides/access/device-login/). ## Create another org [Section titled “Create another org”](#create-another-org) ```bash nodus create org acme --display-name "Acme Research" ``` You become its Owner. Every org starts with a project named `default`. ## Switch orgs [Section titled “Switch orgs”](#switch-orgs) The console has an org switcher. On the CLI, each org is a context: ```bash nodus config get-contexts nodus config use-context acme nodus get jobs --org acme # one command in another org ``` API calls choose an org with the `Nodus-Org` header (the org name or id). Without it, a signed-in user acts in the org they used last. ## Org settings and activity [Section titled “Org settings and activity”](#org-settings-and-activity) Admins can rename the org; its name (the slug in console URLs) stays the same: ```bash nodus patch org acme --patch '{"displayName": "Acme Labs"}' nodus get org acme -o yaml # role, seats used and your org's enabled features ``` Admins and Owners see the org’s activity under **Settings → Activity**: sign-ins, key, member and invite changes, each with who did it and when, for the last day up to the last year, and **Export CSV** downloads what you see. The API serves it at `GET /apis/nodus.dev/v1/auditevents`, oldest first: `?action=auth.` selects sign-ins only, `since` sets the start, and a full page carries `metadata.continue` to pass as `continue` for the next one. # SSH keys > Add the public keys you use to open a shell in a Workspace. An **SSH key** belongs to you, not to an org. Add it once and it works in the Workspaces of every org you belong to. Console › Account › SSH keys lists yours whatever org is selected. ```bash nodus create sshkey laptop --from-file ~/.ssh/id_ed25519.pub nodus get sshkeys ``` Without a key yet, let the CLI make one. `--generate` writes a new Ed25519 key pair to `~/.ssh/id_ed25519` (the public key next to it, with `.pub`) and adds the public key. The private key file is readable by you alone, and `nodus` never sends it anywhere. It refuses to overwrite a file that is there: use `--key-file PATH` to write elsewhere. Add a passphrase afterwards with `ssh-keygen -p -f ~/.ssh/id_ed25519`. ```bash nodus create sshkey --generate ``` Leave the name out and the key is called `-`. `--from-file -` reads the key from standard input, and a file that holds a private key is refused. Or over the API: ```bash curl -X POST "$NODUS_API_URL/apis/nodus.dev/v1/sshkeys" \ -H "Authorization: Bearer $NODUS_API_KEY" -H "Content-Type: application/json" \ -d "{\"name\": \"laptop\", \"publicKey\": \"$(cat ~/.ssh/id_ed25519.pub)\"}" ``` Nodus accepts `ssh-ed25519`, `ecdsa-sha2-*` and `ssh-rsa` keys of at least 3072 bits. It stores the key without its comment and shows its `SHA256:` fingerprint, which is what `ssh-keygen -lf ~/.ssh/id_ed25519.pub` prints. Each name and each key can be added once, and you can hold up to 20 keys. ## Who can connect [Section titled “Who can connect”](#who-can-connect) `nodus ssh workspace/` authenticates with your keys when you are an Owner, Admin or Member of the Workspace’s org and your membership covers its project. Viewers cannot open a shell. When you leave an org, your keys stop working in its Workspaces, and they keep working everywhere else. ## Removing a key [Section titled “Removing a key”](#removing-a-key) ```bash nodus delete sshkey laptop ``` Deleting a key ends the live sessions opened with it, and the next login with it is refused. # Agents > Run an agent on Claude in its own sandbox, send it follow-up prompts, and pay for the model from your credits. An **Agent** is a definition: a system prompt, the Claude access it runs on and a cost cap. An **AgentRun** is one conversation with it. Each run works in its own sandbox, where the model runs shell commands to check its work, and it survives restarts of Nodus by replaying recorded results. An interrupted command or model call using your own key parks with `NeedsResolution` when its outcome is uncertain; inspect its effects before starting replacement work. Agents, parallel AgentGroups and Environment evaluation batches are in Beta. Each run is one agent conversation. ## Sign in and create an agent [Section titled “Sign in and create an agent”](#sign-in-and-create-an-agent) agent.yaml ```yaml apiVersion: nodus.dev/v1 kind: Agent metadata: name: hello-agent spec: system: | You are a careful assistant. Run shell commands in your sandbox to check your work, and answer in one short paragraph. # Nodus's Claude: each prompt is routed to Haiku, Sonnet, Opus or Fable, and billed from your credits. model: access: Nodus perRunMaxCostUSD: "0.50" ``` ```bash nodus login nodus apply -f agent.yaml ``` The agent uses Nodus’s Claude. Every prompt goes to a small router that picks the right Claude model for it (Haiku for simple requests up to Fable for demanding ones). That model answers the whole prompt, including its commands, so the model never changes in the middle of a turn. If the router is unavailable, Sonnet answers. Restrict the choice with `spec.model.families`, for example `[haiku, sonnet]`. ## Start a run [Section titled “Start a run”](#start-a-run) run.yaml ```yaml apiVersion: nodus.dev/v1 kind: AgentRun metadata: name: hello-agent-run spec: agent: hello-agent input: "Use the shell to print the Python version, then tell me which one it is." ``` ```bash nodus apply -f run.yaml nodus wait agentrun/hello-agent-run --for=jsonpath='{.status.phase}'=Succeeded --timeout=8m nodus get agentrun/hello-agent-run -o yaml ``` Or without a file: `nodus create agentrun triage --agent hello-agent --prompt "Summarize the logs"`. To start from Nodus’s ready-made assistant, name the template `claude-assistant`. `status.turns` lists each prompt with the model that served it, how it was chosen (`jev` or `fallback`) and its cost. `status.answer.preview` holds the first 4 KiB of the final answer; `nodus agentrun answer` prints all of it. ```bash nodus agentrun answer hello-agent-run nodus agentrun steps hello-agent-run ``` `steps` lists what the run did in order: creating its sandbox, each routing and model call, each command. A step that shows `Completed` is never repeated, even after a restart. The console’s Agents page shows the same: each run with the model of its latest turn and its cost, and on a run every turn with the model that served it, how that model was chosen and what the turn cost, then the answer and the recorded steps. ## From Python [Section titled “From Python”](#from-python) ```python import nodus agent = nodus.ClaudeAgent("helper", system="You are careful.", families=["haiku", "sonnet"]) print(agent.remote("Use the shell to print the Python version")) # runs to the end, returns the answer run = agent.submit("Summarize the logs", keep_alive=True) # a handle: run.answer(), run.steps(), run.cancel() run.send("message", "Now the staging logs") ``` `ClaudeAgent` creates the agent on first use. Pass `api_key_secret="anthropic-key"` to use your own Anthropic key, or `ClaudeAgent.from_name("claude-assistant", project="nodus")` to run the ready-made template. To run many at once from Python, see [Run agents in parallel](/docs/guides/python/agents/#run-agents-in-parallel). ## Keep a run open for follow-ups [Section titled “Keep a run open for follow-ups”](#keep-a-run-open-for-follow-ups) Set `spec.keepAlive: true` and the run waits for messages after each turn instead of finishing. While it waits it uses no model and keeps its sandbox, which can still accrue compute charges. Set `spec.deadline` to bound the run’s lifetime. The deadline also applies while waiting for a message: the run fails with `DeadlineExceeded` and its sandbox is deleted without another model call. ```bash nodus apply -f run.yaml # with keepAlive: true nodus agentrun send hello-agent-run "Now do the same for the staging logs" --key staging-1 ``` Sending the same `--key` again delivers the message once. A run that is still working reads the message after its current turn; a finished run refuses it with `AgentRunFinished`. Cancel a run with `nodus cancel agentrun/NAME`; its sandbox is deleted. The same calls are REST: `POST …/agentruns/NAME/messages` with `{"payload": "…", "messageKey": "…"}`, and `GET` on `…/steps` and `…/answer`. ## Run agents in parallel [Section titled “Run agents in parallel”](#run-agents-in-parallel) An **AgentGroup** runs many runs of one agent at once. You choose how many run at the same time and how much the whole group may spend. The other runs wait their turn, and one command cancels or deletes all of them. Create the agent and the group, then one run for each task: agent.yaml ```yaml apiVersion: nodus.dev/v1 kind: Agent metadata: name: parallel-agents-worker spec: system: | Follow the task's requested answer format; otherwise answer in one short sentence. # Nodus's Claude, limited to Haiku so each run stays far below its cap. model: access: Nodus families: [haiku] maxTokens: 1024 perRunMaxCostUSD: "0.10" ``` group.yaml ```yaml apiVersion: nodus.dev/v1 kind: AgentGroup metadata: name: parallel-agents spec: agent: parallel-agents-worker limits: maxActive: 2 # at most two runs go at once maxCostUSD: "0.60" # a run starts only while its own cap still fits under what is left ``` runs.yaml ```yaml apiVersion: nodus.dev/v1 kind: AgentRun metadata: name: parallel-agents-a spec: agent: parallel-agents-worker group: parallel-agents taskKey: a input: "What is the capital of France?" --- apiVersion: nodus.dev/v1 kind: AgentRun metadata: name: parallel-agents-b spec: agent: parallel-agents-worker group: parallel-agents taskKey: b input: "What is the capital of Japan?" --- apiVersion: nodus.dev/v1 kind: AgentRun metadata: name: parallel-agents-c spec: agent: parallel-agents-worker group: parallel-agents taskKey: c dependsOn: [a, b] # c waits until a and b have succeeded input: "Say in one sentence that both questions have been answered." ``` ```bash nodus apply -f agent.yaml nodus apply -f group.yaml nodus apply -f runs.yaml nodus patch ag/parallel-agents --patch '{"spec":{"sealed":true}}' ``` Or without a file: `nodus create agentgroup parallel-agents --agent parallel-agents-worker --max-active 2 --max-cost 0.60`. Each run is a task of the group: `spec.group` names the group and `spec.taskKey` names the task (a DNS label that is unique in the group). The run is named `-`; leave `metadata.name` out or set exactly that. The run uses the group’s agent, so `spec.agent` must match it. The last command seals the group, which says that no more runs are coming (see below). ### Watch progress and read the answers [Section titled “Watch progress and read the answers”](#watch-progress-and-read-the-answers) ```bash nodus wait ag/parallel-agents --for=jsonpath='{.status.phase}'=Succeeded --timeout=12m nodus get ag nodus get ar --field-selector spec.group=parallel-agents nodus agentrun answer parallel-agents-c ``` `nodus get ag` lists each group with its phase, `RUNS` (finished successfully out of all) and `ACTIVE` runs, and its cost. `status.counts` splits the runs into `queued`, `active`, `waiting` (for a message or for funds), `succeeded`, `failed` and `cancelled`. `status.blockedReasons` says why queued runs have not started and how many wait for each reason: `DependencyWait`, `MaxActiveReached` or `MaxCostReached`. A queued run’s own `status.reason` is `DependencyWait`, or `GroupWait` while it waits for a free slot or for room under the cost cap. The runs you list are ordinary runs, so `status.answer`, `nodus agentrun answer` and `nodus agentrun steps` work on each of them. A group is `Running` until it is sealed and every run has finished. It then becomes `Succeeded`, or `Failed` with reason `RunsFailed` when a run failed or was cancelled because a dependency failed. A sealed group with no runs succeeds. ### How many run at once [Section titled “How many run at once”](#how-many-run-at-once) `spec.limits.maxActive` is how many runs the group lets run at the same time (default 10, 1 to 1000). Change it on a live group: ```bash nodus patch ag/parallel-agents --patch '{"spec":{"limits":{"maxActive":4}}}' ``` `spec.limits.maxPending` is how many unfinished runs the group takes in (default 10000); a run beyond it is refused. ### Dependencies [Section titled “Dependencies”](#dependencies) `spec.dependsOn` lists the task keys a run waits for. It starts only after every one of them has succeeded, and until then it is queued with reason `DependencyWait`. A dependency must already be in the group, so create the runs it depends on first (earlier in the same file works). When a dependency fails, the run is cancelled with reason `DependencyFailed`; when a dependency was cancelled, the run is cancelled with reason `DependencyCancelled`. ### The group cost cap [Section titled “The group cost cap”](#the-group-cost-cap) `spec.maxCostUSD` is one cap for the whole group. A run starts only when its own cap (the agent’s `perRunMaxCostUSD`) still fits under what is left: the cap minus what the group has spent and minus the caps of the runs that have started and not finished. The members together can therefore never spend more than the cap, and runs that do not fit yet wait with `MaxCostReached`. In the example, the $0.60 cap has room for the caps of six runs at $0.10 each before any run has spent anything. Raise the cap on a live group with `nodus patch`; it cannot be lowered. If the cap is smaller than one run’s cap, no run starts: raise it or lower the agent’s `perRunMaxCostUSD`. `status.cost` shows `totalUSD` and `limitUSD`. The cap covers the runs’ Claude usage (model and routing calls), the spend each run’s `perRunMaxCostUSD` limits; the sandbox each run works in is billed on its own and is not counted against either cap. ### Seal, cancel and delete [Section titled “Seal, cancel and delete”](#seal-cancel-and-delete) A group takes new runs until you seal it with `spec.sealed: true`. Sealing cannot be undone, and a sealed group refuses new runs with a conflict. A group finishes only once it is sealed, so seal it after you create the last run. ```bash nodus cancel ag/parallel-agents nodus delete ag/parallel-agents ``` `nodus cancel` cancels every run that has not finished and leaves the finished runs, with their answers, as they are. `status.phase` goes through `Cancelling` to `Cancelled`. `nodus delete` removes the group together with all of its runs. ## Evaluate an agent [Section titled “Evaluate an agent”](#evaluate-an-agent) Create an evaluation group with an Environment at an explicit version. Nodus generates the fixed task batch, runs it through ordinary AgentRuns, and grades completed answers using the Environment’s isolated grader. There is no manual submit or seal step. Set a project Budget to limit total spend: the group cap covers model usage, while agent and grading Sandbox compute is billed separately. After deploying the example’s `parallel-agents-worker` Agent above, run this Python example. It creates two held-out tasks, waits for a terminal result, prints per-case outcomes and deletes the evaluation group afterward. evaluation.py ```python """Evaluate the deployed example agent on a fixed, versioned task batch.""" import nodus def main() -> None: # Model caps exclude Sandbox compute; set a project Budget before running this example. evaluation = nodus.AgentGroup.create( "parallel-agents-eval", agent="parallel-agents-worker", max_active=2, max_cost="0.20", evaluation={ "environment": "nodus/arithmetic-v2@2.0.0", "split": "test", "tasks": 2, "repetitions": 1, "seed": 42, "timeout": "10m", }, ) try: status = evaluation.wait(timeout=660) results = evaluation.results() print( f"{status.phase}: cases={len(results)}, pending={status.evaluation.pending}" ) print("pass rate:", status.evaluation.get("passRate")) print("mean reward:", status.evaluation.get("meanReward")) for result in results: print(result.task_id, result.state, result.get("verdict")) finally: evaluation.delete() if __name__ == "__main__": main() ``` The batch is at most 100 cases including repetitions. Every member pins the same Agent revision; results record the Environment version, resolved image digest, split and seed. Only task prompts reach the agent, never hidden answers. The timeout defaults to 30 minutes, may be 1 minute to 24 hours, and includes manifest preparation, queueing and funding waits from group creation. Each member receives that absolute deadline. Grading also stops at the deadline; completed answers that could not be scored are shown as evaluation timeouts, excluded from pass rate and mean reward. Open the group in the console to see scored cases, pass rate, mean reward, pending cases and execution/grading failures separately. The Results view links each case to its AgentRun and shows grader evidence. A failed run, missing answer or grader failure is unscored and does not become a zero reward. A wrong or invalid answer is a scored result. Per-case records share the member journal’s 30-day payload retention; deleting the group removes its runs and records. Canceling or finishing the group cleans up its grading pool. ## Use your own Anthropic key [Section titled “Use your own Anthropic key”](#use-your-own-anthropic-key) Store the key in a Secret under the key `ANTHROPIC_API_KEY` and point the agent at it. Nodus then charges nothing for the model, and your Anthropic account pays for it. ```yaml spec: model: access: BYOK name: claude-sonnet-5-5 apiKeySecret: anthropic-key ``` ## What a run costs [Section titled “What a run costs”](#what-a-run-costs) On Nodus’s Claude each model call, and the routing call before each turn, is billed from your credits at cost divided by 0.875. `spec.perRunMaxCostUSD` (default $1.00) caps what a run spends on the model. Before each call the run checks that the call, at its largest (the whole conversation as input and `spec.maxTokens` of output), still fits under the cap. When it does not, the run waits (`reason: MaxCostReached`) and makes no more calls. The cap is copied from the agent when the run is created, so changing the agent does not change a run that has started: to continue the work, cancel the run and start a new one. If your credits run out, the run waits instead and continues when you add credits. A run on your own Anthropic key is not charged by Nodus and has no cap to reach. Sandbox compute is billed as for any sandbox. `status.cost` shows the model spend so far. Runs in a group also share the group’s cap (see [Run agents in parallel](#run-agents-in-parallel)). ## Limits [Section titled “Limits”](#limits) * One run works one prompt at a time, with at most `spec.maxToolCalls` commands per turn (default 50). * The input and each message are at most 256 KiB. A run’s recorded payloads are kept 30 days after it ends. * A run created for a stopped agent (`spec.state: Stopped`) is refused with `AgentStopped`. * A group takes at most 10000 unfinished runs. Create runs in batches of up to 100 by posting an `AgentRunList` to `…/agentruns` with an `Idempotency-Key` header; a batch is all-or-nothing, and its runs are created in list order, so list each run after the runs it depends on. See [Durable execution](/docs/concepts/durable-execution/) for what survives a restart. # Ask Nodus > Ask the console assistant why a Job failed, what it cost, or to draft a manifest. It cites the docs, and changes only after you confirm. **Ask Nodus** is the assistant in the console. It answers questions about your org: why a Job restarted, what a Sandbox cost this week, which flags a manifest needs. It uses the same tools as the [MCP server](/docs/guides/mcp/), with your own role and scopes, so it can see and do only what you can. ## Ask a question [Section titled “Ask a question”](#ask-a-question) Open Ask Nodus in the console and type. Some things to try: * “Why did job/train fail?” * “What did my Sandboxes cost this month?” * “How do I make train survive being preempted?” * “Write a manifest that runs `python train.py` on an L4 with a $5 cap.” When an answer relies on the docs, it links the pages, and the links are the pages the assistant actually read. ## Changes need your confirmation [Section titled “Changes need your confirmation”](#changes-need-your-confirmation) The assistant can propose a change, but it cannot make one. When an answer would create, delete, suspend or otherwise change something, the turn stops and shows exactly what will run. For a new object that includes the server’s dry-run: the object as it would be stored, its estimated cost and anything that blocks it. Choose **Confirm** to run it, or **Cancel**. Confirming runs the held change once, with your credential, so the API checks your permissions again. Asking a new question cancels a change you did not confirm. The assistant proposes at most one change per answer. ## Drafts [Section titled “Drafts”](#drafts) When you ask for a manifest, the assistant writes one and checks it with the same dry-run that `nodus apply --dry-run=server` uses. A draft is shown to you only after that check passes. If the check fails, the assistant sees the error and fixes the manifest first. Each draft comes with the equivalent CLI command and Python, and its estimate. A draft is not created until you apply it, or ask the assistant to and confirm. ## Checkpoint suggestions [Section titled “Checkpoint suggestions”](#checkpoint-suggestions) For a Job that would start over after a preemption, the assistant can suggest a manifest that checkpoints, with that manifest’s dry-run. Recovery settings cannot change on a running Job, so the suggestion is a new manifest. Your program still has to write its state to the declared checkpoint paths. ## Your preferences and followed objects [Section titled “Your preferences and followed objects”](#your-preferences-and-followed-objects) You can give the assistant a short note about how you like answers, such as “short answers, show the CLI command”. It reads the note as your preference, and the note cannot change the assistant’s rules. You can also follow up to 8 objects so the assistant knows what you are working on. Both are yours alone in each org. In the API: ```console $ curl -s "$NODUS_API_URL/assistant/v1/profile" -H "Authorization: Bearer $NODUS_API_KEY" $ curl -s -X PUT "$NODUS_API_URL/assistant/v1/profile" -H "Authorization: Bearer $NODUS_API_KEY" \ -H "Content-Type: application/json" -d '{"note": "Short answers.", "revision": 0}' ``` The note is at most 8 KiB. Send back the `revision` you read, and the update fails with a conflict if someone else changed the note in between. `/assistant/v1/watches` takes `{"watches": ["", …]}`. ## Chats and privacy [Section titled “Chats and privacy”](#chats-and-privacy) Your chats are stored per org, for you only. A chat holds up to 200 questions. Each answer is stored with a checksum, and deleting a chat deletes its answers. A tool the assistant ran is recorded by name with checksums of its arguments and its result, never their content, so file contents and command output are not kept. Chats are kept at most 400 days. Nodus chooses a Claude model for each question and uses it for the whole answer. Nodus pays for the model calls; they do not use your credits. ## Limits [Section titled “Limits”](#limits) | Limit | Value | | ---------------------- | --------------------------------------------------------------------------- | | Questions | 1 per second per user, with bursts of up to 20 | | Steps for one question | 12 model rounds and 32 tool calls | | Time for one question | 180 seconds | | Drafts | 1 draft request every 5 seconds per user | | Daily allowance | Each org has a daily allowance for the assistant, which resets at 00:00 UTC | When a limit is reached the assistant says so and what happens next, instead of failing. Your Jobs, the docs and the CLI are not affected. Note The assistant reads logs and files as data, never as instructions. Text inside a log cannot make it run anything, and anything it proposes still needs your confirmation. # Usage and billing > Add credits, redeem a code, cap spending with Budgets, and see exactly what every run cost. Nodus is prepaid: you add credits, and every run reserves funds before it starts. New orgs get [starter credit](/docs/guides/billing/credits/) to try things out. This guide covers the everyday tasks; the [billing concept page](/docs/concepts/billing/) explains holds, captures and limits. ## Check your balance [Section titled “Check your balance”](#check-your-balance) ```console $ nodus billing Available $17.42 purchased $15.80 · credits $3.52 (starter, expires Oct 31) Reserved $1.90 job/train-a, sandbox/sb-3, function/embed This month $232.58 budget research-monthly 42 % of $500.00 Auto-recharge on add $50.00 below $10.00 · visa •••• 4242 ``` In the console, open **Usage & billing**. The header shows your available balance everywhere, in amber when it is low. ## Add credits [Section titled “Add credits”](#add-credits) ```console $ nodus billing top-up 20 Opening https://checkout.stripe.com/c/pay/cs_live_... (expires in 60 min) Waiting for payment... added $20.00. Available $37.42. Receipt: https://invoice.stripe.com/i/... ``` Top-ups are between $5.00 and $1,000.00 in whole cents, up to $5,000 per org per day. Payment happens on a Stripe-hosted page; Nodus never sees your card number. In the console, **Add credits** offers $10, $20, $50, $100 or a custom amount and brings you back to Billing when the payment completes. A top-up settles any unpaid charges first. Every top-up has a receipt and an invoice PDF under **Receipts & invoices** and in `nodus billing receipts`. Topping up needs the `billing:write` permission, which org Owners and Admins have. ## Auto-recharge [Section titled “Auto-recharge”](#auto-recharge) Auto-recharge adds credit when your available balance drops below a threshold, so long runs never stop for money: ```console $ nodus billing auto-recharge --threshold 10 --amount 50 ``` The first time, this opens a card setup page. Each recharge is a Stripe invoice with its own receipt. At most 5 recharges or $5,000 run per day, and three failed charges in a row turn auto-recharge off and email you. Turn it off with `nodus billing auto-recharge --off`. **Save** updates the amounts and warning level without turning auto-recharge on or off. Use **Turn on** or **Turn off** to change that setting. If another admin or failed payments change it while you edit, the console refreshes the settings and asks you to review them before saving again. ## Redeem a promo code [Section titled “Redeem a promo code”](#redeem-a-promo-code) ```console $ nodus billing redeem LAUNCH25 Redeemed ****CH25: $25.00 of credit, expires 2026-11-30. ``` See [Promo codes](/docs/guides/billing/promo-codes/) for limits and errors. ## See what you spent [Section titled “See what you spent”](#see-what-you-spent) ```console $ nodus get usage --group-by project,label:owner --since 30d PROJECT OWNER AMOUNT research ml $212.41 default - $20.17 $ nodus get usage --group-by project,meter --since 30d -o csv > usage.csv $ nodus get transactions --since 7d ``` Group usage by `project`, `kind`, `label:`, `meter`, `day`, `segment` or `rank`. `segment` splits machine time into `Boot`, `Restore`, `Running` and `Teardown`, and `rank` itemizes the members of a multi-node run. The console **Usage** tab shows the same data as a daily chart and a table, and every table exports CSV. Transactions list every change to your balance, with the balance after it: top-ups, grants, captures, storage, egress, refunds and adjustments. [What you pay for](/docs/guides/billing/what-you-pay-for/) lists every kind of time and who pays for it. ## Cap spending with a Budget [Section titled “Cap spending with a Budget”](#cap-spending-with-a-budget) A Budget is an enforced limit over the org, a project or a label selector, per month or in total. When a `Block` Budget is exhausted, work in its scope stops gracefully and new work is refused until the next period: budget.yaml ```yaml apiVersion: nodus.dev/v1 kind: Budget metadata: name: examples-billing-monthly spec: limitUSD: "25.00" period: Monthly scope: project: examples action: Block thresholds: [50, 80, 100] ``` ```console $ nodus apply -f budget.yaml $ nodus get budgets ``` You get an email at each threshold. The console **Budgets** tab creates and edits Budgets and previews which objects a scope matches now. To cap one person, scope a Budget to them by email. It counts everything they start, with any of their API keys, including inference requests: ```console $ nodus create budget ada-monthly --limit 300 --period Monthly --scope-member ada@example.com ``` ## When work is refused for money [Section titled “When work is refused for money”](#when-work-is-refused-for-money) A create that cannot be funded fails with `402` and the exact amounts, and the console shows an **Add credits** action: | Error | What to do | | --------------------- | --------------------------------------------------------------- | | `InsufficientCredits` | Add credits, or lower `spec.maxCostUSD` | | `BudgetExceeded` | Raise the Budget’s `spec.limitUSD`, or wait for its next period | | `ArrearsOutstanding` | Add credits; the top-up settles the unpaid charges first | | `PaymentDisputed` | Contact support; new work waits until the dispute closes | ## Payment methods and billing details [Section titled “Payment methods and billing details”](#payment-methods-and-billing-details) `nodus billing portal` opens the Stripe Customer Portal, where you update cards, your billing email, address and tax ID, and download past invoices. ## Learn more [Section titled “Learn more”](#learn-more) * [Credits](/docs/guides/billing/credits/): the starter credit, grants and expiry * [Promo codes](/docs/guides/billing/promo-codes/) * [Refunds](/docs/guides/billing/refunds/) * [What you pay for](/docs/guides/billing/what-you-pay-for/) # Add a card before running work > Save a payment method without buying credits. Open **Billing → Payment method → Add a card**, or run: ```sh nodus billing portal --setup ``` Stripe saves your card without charging it or enabling auto-recharge. Your $30 signup credit keeps its original 30-day expiration. New work needs a verified card even when promotional credit covers its entire cost. An organization billing administrator must add the card. Other members can view its status. After returning from Stripe, wait for Billing to show that the card is ready. Returning to the page alone does not complete verification. If you canceled setup, choose **Add a card** again. CLI and SDK requests receive `PaymentMethodRequired` until verification completes. `PaymentVerificationUnavailable` means to retry shortly. Removing or letting the card expire blocks new work. Existing running work continues within its available credits and budgets. A restart or new allocation checks the payment method again. Adding a card does not let work exceed your prepaid balance. # Auto-recharge > Top up automatically from a saved card when your balance falls below a threshold, so running work never stops for money. Auto-recharge buys credits for you when your available balance falls below a threshold you choose. Each recharge is charged to a saved card and emailed to you as a receipt, exactly like a top-up you buy yourself. It is off until you turn it on. ## Turn it on [Section titled “Turn it on”](#turn-it-on) ```bash nodus billing auto-recharge --threshold 10 --amount 50 ``` This recharges $50 whenever your available balance drops below $10. If no card is on file yet, the command first opens a Stripe page to save one, then turns auto-recharge on. You can also save a card while buying credits with `nodus billing top-up 20 --save-card`, and the console billing page has an auto-recharge card that does both. The same settings live on your org’s `BillingAccount`: ```yaml apiVersion: nodus.dev/v1 kind: BillingAccount metadata: name: default spec: autoRecharge: enabled: true thresholdUSD: "10.00" # at least 5.00 amountUSD: "50.00" # 5.00 to 1000.00 ``` | Setting | Default | Allowed | | -------------- | --------- | ------------------ | | `thresholdUSD` | `"10.00"` | $5.00 or more | | `amountUSD` | `"20.00"` | $5.00 to $1,000.00 | Turning it on without a saved card is refused; check `PaymentMethodPresent` in `nodus billing`, or save a card through the billing portal (`nodus billing portal`). ## When it recharges [Section titled “When it recharges”](#when-it-recharges) Nodus checks your balance each time it renews the hold of running work, settles inference or storage, or refuses a new hold. It recharges when all of these are true: * your available balance is below `thresholdUSD`; * a card is on file; * no other auto-recharge is still being paid; * the org has had fewer than 5 auto-recharges, and less than $5,000 of them, today (UTC); * the last 3 auto-recharges did not all fail. A recharge usually lands in seconds. Work whose funding is running short keeps running while it is paid, so a successful recharge within that window means nothing stops. ## See what it did [Section titled “See what it did”](#see-what-it-did) Each recharge is a `TopUp` with `spec.origin: AutoRecharge`: ```bash nodus get topups -o wide nodus billing receipts ``` `nodus billing` shows the auto-recharge state from `status.autoRecharge`: `consecutiveFailures`, `lastAttemptTime`, `lastResult` and, when it has been switched off, `disabledReason`. ## When a recharge fails [Section titled “When a recharge fails”](#when-a-recharge-fails) If the card is declined, the recharge `TopUp` ends `Failed` (`PaymentFailed`), nothing is credited, and the billing email and the org’s owners and admins get an email. Auto-recharge tries again the next time the balance is checked. After 3 failures in a row, Nodus turns auto-recharge off, sets `status.autoRecharge.disabledReason` to `ConsecutiveFailures`, shows a banner in the console and emails the same people. To turn it back on: 1. Update the card in the billing portal: `nodus billing portal`. 2. Run `nodus billing auto-recharge --threshold 10 --amount 50` again, or set `enabled: true`. If your balance runs out while auto-recharge is off, work stops gracefully as described in [When money runs out](../when-money-runs-out/). ## Turn it off [Section titled “Turn it off”](#turn-it-off) ```bash nodus billing auto-recharge --off ``` The saved card stays on file for later; remove it in the billing portal. # Budgets and spending caps > Put an enforced limit on what an org, a project or a labelled group of work can spend, and cap any single run. Nodus enforces two kinds of spending limits. Both are checked before any paid work starts and again every few minutes while it runs, so a limit always stops spend instead of reporting it afterwards. * A **Budget** limits what matching work across your org can spend in a month or in total. * **`maxCostUSD`** caps one object: a Job, Pipeline, Sweep, TrainingJob, Sandbox, Workspace, Function, Agent or InferenceEndpoint. A cap on a Pipeline or Sweep is shared by every run it creates. ## Create a Budget [Section titled “Create a Budget”](#create-a-budget) ```yaml apiVersion: nodus.dev/v1 kind: Budget metadata: name: research-monthly spec: limitUSD: "500.00" period: Monthly # a UTC calendar month; Total counts from creation scope: project: research # leave out for the whole org selector: matchLabels: {team: nlp} action: Block # Notify only sends notices thresholds: [50, 80, 100] notify: emails: [research-leads@example.com] ``` ```bash nodus apply -f budget.yaml nodus get budget research-monthly ``` `nodus get budget` shows what is spent, what is held for running work, what remains and the thresholds already notified this period. An org can have up to 50 Budgets. ## What a Budget covers [Section titled “What a Budget covers”](#what-a-budget-covers) A Budget counts every hold and charge of work in its project whose labels matched the selector **when the work was created**. Changing a running Job’s labels does not move it out of a Budget. A Budget you create while work is already running starts covering that work within 5 minutes. ## What happens at the limit [Section titled “What happens at the limit”](#what-happens-at-the-limit) With `action: Block`: * New work in scope is refused with `402 BudgetExceeded`, naming the Budget, what is left and what the work needs: ```text budget "research-monthly" has $0.40 left of $500.00 this month; job "train-a" needs $3.20. fix: raise spec.limitUSD on budget/research-monthly, or wait until 2026-11-01 ``` * Running work in scope stops gracefully, as described in [When money runs out](../when-money-runs-out/), with `Funded=False` and reason `BudgetExceeded`. * Raising `spec.limitUSD`, switching to `Notify`, deleting the Budget or the start of the next month resumes it. With `action: Notify`, nothing is refused; you only get the notices. ## Notices [Section titled “Notices”](#notices) Each threshold is sent once per period as a `BudgetThreshold` Event, a `budget.threshold` webhook and an email to the billing email, org owners and admins, and the addresses in `spec.notify.emails`. Reaching a `Block` limit also sends `BudgetExceeded` (webhook `budget.exceeded`). ## Cap a single run [Section titled “Cap a single run”](#cap-a-single-run) ```yaml kind: Job spec: maxCostUSD: "25.00" ``` A Job never holds more than its cap. When its spend reaches the cap, it checkpoints and becomes `Suspended` with reason `MaxCostReached`; raising `maxCostUSD` resumes it from the checkpoint. You can raise a cap at any time, but not lower it. A Pipeline, Sweep or TrainingJob shares its cap with the Jobs it creates. Each child must fit both its own cap and its parent’s remaining cap, along with every matching Budget. A dry-run reports an insufficient spending limit in `status.estimate.blockingReasons` without reserving funds. # Buy credits > Add prepaid credits with a card through Stripe Checkout from the console, the CLI or the API, and add the $20 Indra plan. Nodus is prepaid. You buy credits, and paid work draws on them through a funded hold placed before it starts. You pay on a Stripe-hosted page, so Nodus never sees your card number. Every purchase is a `TopUp` object that you can list and follow until the credit is on your balance. ## Buy from the console [Section titled “Buy from the console”](#buy-from-the-console) Open **Usage & billing**, choose **Add credits**, and pick $10, $20, $50, $100 or a custom amount. The console sends you to Stripe Checkout and brings you back to the billing page, which shows the top-up until it is paid. ## Buy from the CLI [Section titled “Buy from the CLI”](#buy-from-the-cli) ```bash nodus billing top-up 20 ``` The CLI opens the payment page in your browser, waits until the payment is credited and prints your new balance. | Flag | Does | | ------------- | ------------------------------------------------------------------------ | | `--save-card` | Saves the card for [auto-recharge](../auto-recharge/) | | `--no-open` | Prints the payment link instead of opening a browser, for a remote shell | | `--no-wait` | Returns as soon as the payment link exists | ## Buy through the API [Section titled “Buy through the API”](#buy-through-the-api) A top-up is an ordinary resource. Create it, then open its `status.checkoutURL`, which appears within about a second: ```yaml apiVersion: nodus.dev/v1 kind: TopUp metadata: name: october-credits spec: amountUSD: "50.00" savePaymentMethod: false # true also saves the card for auto-recharge ``` ```bash nodus apply -f topup.yaml nodus get topup october-credits -o yaml ``` The REST path is `/apis/nodus.dev/v1/topups`. By default Checkout returns to the console billing page; set `spec.successURL` and `spec.cancelURL` to return somewhere else on the console, or to an `http://127.0.0.1` address for a local tool. A top-up cannot be edited or deleted; an unpaid one expires by itself. ## Limits [Section titled “Limits”](#limits) | Limit | Value | | ----------------------- | -------------------------------------- | | Amount of one top-up | $5.00 to $1,000.00, in whole cents | | Top-ups per org per day | $5,000 in total, per UTC day | | Payment page | Payable for 1 hour after it is created | An amount outside the range is refused with `422 Invalid` before anything is charged. ## Follow a top-up [Section titled “Follow a top-up”](#follow-a-top-up) ```bash nodus get topups -o wide ``` | Phase | Reason | Meaning | | ----------- | ---------------- | ---------------------------------------------------------------------------------------------------- | | `Queued` | | Nodus is creating the payment page | | `Running` | | The payment page is open and waiting for you | | `Running` | `PaymentPending` | You paid with a method that confirms later, such as a bank debit | | `Succeeded` | | Paid and credited; `status.creditedUSD` is on your balance and `status.receiptURL` links the receipt | | `Failed` | `Expired` | Nobody paid within the hour; nothing was charged | | `Failed` | `PaymentFailed` | The payment did not complete; nothing was credited | A declined card does not fail the top-up: the payment page stays open so you can try another card. When a payment fails after you submit it, such as a bank debit that does not clear, Nodus emails your billing email. ## When the credit arrives [Section titled “When the credit arrives”](#when-the-credit-arrives) Nodus credits each payment once, as soon as Stripe confirms it, which usually takes a few seconds. If a confirmation from Stripe is delayed or lost, Nodus asks Stripe about the payment itself within the hour and again in a nightly check, so a paid top-up is always credited, and never twice. If your org owes arrears, the top-up pays them first and the rest goes to your balance. Work that stopped because money ran out resumes by itself; see [When money runs out](../when-money-runs-out/). Stripe emails a receipt for every paid top-up to your billing email. See [Receipts and invoices](../receipts/). ## The Indra plan [Section titled “The Indra plan”](#the-indra-plan) The Indra plan costs $20 a month. Each paid month adds a $20 credit that only Indra (`nodus/indra`) requests can spend, at list price. Indra requests use that credit before your other credits, and everything else draws on your regular balance. Unused plan credit expires when the month it belongs to ends; it does not carry over. Start the plan by opening an Indra plan checkout session (type `ComposerSubscription`) and paying on the page it returns: ```bash curl -X POST "https://api.nodus-compute.ai/apis/nodus.dev/v1/billingaccounts/default/sessions" \ -H "Authorization: Bearer $NODUS_TOKEN" -H "Idempotency-Key: $(uuidgen)" \ -d '{"type": "ComposerSubscription"}' ``` The response is `{url, expirationTime}`. An org can have one Indra plan; a second request while the plan is active is refused with `409 Conflict`. * **Renewal.** Stripe charges the card each month and emails the invoice. A new month’s credit is added only once its invoice is paid; while a renewal payment is failing, Stripe retries it and Indra requests draw on your regular credits. * **Cancel.** Open the billing portal with `nodus billing portal` and cancel the plan there. It ends at the end of the paid month: that month’s credit stays usable until then, and nothing further is charged or added. # Credits > The starter credit, credit grants, the order credits are spent in, and when they expire. Credit grants are credit Nodus gives you, as opposed to credit you buy. Each grant is a `CreditGrant` you can read but not change: ```console $ nodus get creditgrants NAME SOURCE AMOUNT REMAINING EXPIRES grt_01j9x... Starter $30.00 $12.40 2026-10-31 grt_01j9y... Promo $25.00 $25.00 2026-11-30 ``` The console lists them under **Usage & billing → Credits**, with what each has spent and when it expires. ## Starter credit [Section titled “Starter credit”](#starter-credit) The first org you create gets **$30.00** of starter credit once your email address is verified. It expires 30 days after the org is created. Each person gets starter credit once, however many orgs they create. Orgs that only have starter credit keep tighter limits until their first top-up: 20 GiB of egress a day, $1 a day of console assistant use and no multi-node runs. ## Where grants come from [Section titled “Where grants come from”](#where-grants-come-from) | Source | How you get it | | ---------- | ------------------------------------------------------------------------ | | `Starter` | Your first org, once your email is verified | | `Promo` | [Redeeming a promo code](/docs/guides/billing/promo-codes/) | | `Admin` | Credit added by Nodus support, for example a pilot or a migrated balance | | `Goodwill` | Credit from Nodus support for a problem on our side | Some plan grants pay only for Indra (`nodus/indra`) model calls; they show `nodus/auto`, Indra’s other name, as their scope and are never spent on anything else. ## Spending order [Section titled “Spending order”](#spending-order) Charges draw from grants first, soonest expiry first, then from grants that never expire, and only then from purchased credit. You never have to pick which credit pays. ## Expiry [Section titled “Expiry”](#expiry) You get an email 7 days and 1 day before a grant expires. At its expiry the unspent remainder leaves your balance at once, even if a run is still using it; the run keeps going on its other credit. Grants are never refunded or turned into cash. A grant also settles unpaid charges first, the same way a top-up does. # Partner offers > Apply an eligible partner offer and understand promotional matching. Open **Billing → Partner offers**, enter your code and optionally add your company or batch. An organization admin submits the request; a platform administrator verifies eligibility before credit is granted. Organization members can view progress. For the YC offer, promotional credit reaches **$100 total**, including earlier grants even if spent or expired. A previous $30 starter grant therefore leaves **$70 extra**. You keep the existing **10 GB storage inclusion**. After approval, earn **10% back on the first $10,000 of eligible usage**, for a maximum **$1,000 bonus**. Only captured usage paid from purchased credit qualifies. Usage paid by starter credit, promotional grants or earned bonuses does not. Earned credit does not expire. Usage corrections adjust the matching benefit, including after an offer is revoked. You can also use the CLI: ```sh nodus billing deal nodus billing deal YC --note 'Company and batch' ``` A request snapshots the terms you applied for. Later offer edits do not change them. Approval keeps the usual card prerequisite and usage prices. If an offer is revoked, future matching stops and existing earned credit remains, subject to corrections. # Promo codes > Redeem a promo code for credit, and what each redemption error means. A promo code adds a [credit grant](/docs/guides/billing/credits/) to your org. Redeem it from the CLI or from **Usage & billing → Credits → Redeem code** in the console: ```console $ nodus billing redeem LAUNCH25 Redeemed ****CH25: $25.00 of credit, expires 2026-11-30. ``` Codes are not case-sensitive, and spaces around them are ignored. Redeeming needs the `billing:write` permission. The response shows the code masked to its last four characters; Nodus stores only a keyed digest of each code. ## Errors [Section titled “Errors”](#errors) | Error | Meaning | | --------------------- | ---------------------------------------------------------------------------- | | `404 NotFound` | The code does not exist, or it has been withdrawn | | `410 Expired` | The code’s last redemption date has passed | | `409 Conflict` | Your org has already redeemed this code, or the code has been fully redeemed | | `429 TooManyRequests` | More than 5 attempts in a minute; wait a minute and try again | Retrying a redemption with the same `Idempotency-Key` returns the first result and never adds credit twice. The CLI and console set the key for you. ## Expiry [Section titled “Expiry”](#expiry) A promo grant expires on the date the code sets, counted from when you redeem it. Codes without one never expire. # Receipts and invoices > Find the receipt and invoice for every top-up and plan payment, change where receipts are sent, and manage your card and billing details. Every payment you make to Nodus, whether a top-up you buy, an auto-recharge or an Indra plan month, has a Stripe-hosted invoice with a PDF, and Stripe emails a receipt for it. Nodus does not send postpaid bills: credits are prepaid, and what they paid for is in your [usage records](../usage-and-costs/). ## Where receipts go [Section titled “Where receipts go”](#where-receipts-go) Receipts go to the billing email of your org’s `BillingAccount`, which defaults to the email of the owner who created the org. To send them somewhere else, such as a finance mailbox, change `spec.billingEmail`: ```yaml apiVersion: nodus.dev/v1 kind: BillingAccount metadata: name: default spec: billingEmail: finance@example.com ``` ```bash nodus apply -f billing.yaml ``` The new address applies to the next payment and is also where Nodus sends money notices, such as a failed auto-recharge, together with the org’s owners and admins. ## Find a receipt [Section titled “Find a receipt”](#find-a-receipt) ```bash nodus billing receipts ``` This lists your top-ups, Checkout and auto-recharge alike, with a link to each invoice page and its PDF. The same links are on each `TopUp`: ```bash nodus get topup october-credits -o yaml ``` | Field | Links to | | ---------------------- | ------------------------------------------------------------------------- | | `status.receiptURL` | The Stripe-hosted invoice page, which shows the payment and the card used | | `status.invoicePDFURL` | The invoice as a PDF | The links appear a few seconds after the payment succeeds, once Stripe has finalized the invoice. ## The billing portal [Section titled “The billing portal”](#the-billing-portal) ```bash nodus billing portal ``` This opens the Stripe billing portal for your org, where you can: * see and download every invoice and receipt, including Indra plan invoices; * add, replace or remove the card used for auto-recharge and the Indra plan; * change the name, address and other details printed on your invoices; * cancel the Indra plan. The portal link is valid for 5 minutes; run the command again for a new one. From the API, create a session with `POST /apis/nodus.dev/v1/billingaccounts/default/sessions` and `{"type": "Portal"}`. Nodus does not charge sales tax or VAT on top-ups. ## Refunds [Section titled “Refunds”](#refunds) Refunds cover unspent purchased credit and are made by Nodus support. A refunded top-up shows the amount in `status.refundedUSD`, the credit leaves your balance when the refund is approved, and the refund appears on the top-up’s invoice page. Credit from grants and promo codes is not refundable. ## Questions about a charge [Section titled “Questions about a charge”](#questions-about-a-charge) Contact Nodus support before disputing a charge with your bank. While a dispute is open, new work that needs credit is refused with `402 PaymentDisputed`; running work is not stopped. When the dispute closes in your favor the credit returns to your balance. # Refunds > How to get unspent purchased credit back to your card, and what can be refunded. You can get **unspent purchased credit** back to the card that paid for it. Credit grants, including starter credit and promo codes, are never refunded. ## Ask for a refund [Section titled “Ask for a refund”](#ask-for-a-refund) Contact support with your org and the top-up you want refunded. Support checks the request and starts the refund; refunds above $500 are approved by a second person at Nodus. You do not need to stop your work first. The amount leaves your available balance as soon as the refund is approved, so it cannot be spent while the card refund is processing. The top-up shows the refund under **Usage & billing → Receipts & invoices**, and `nodus get transactions` shows a `RefundRequest` and then a `Refund`. Your card is usually credited within 5 to 10 business days. If the card refund fails, a `RefundFailed` transaction returns the amount to your balance. ## What can be refunded [Section titled “What can be refunded”](#what-can-be-refunded) * At most what you paid on that top-up, less anything already refunded on it. * At most your unspent purchased credit at approval: credit already used for runs is not refunded. * One refund per top-up at a time, in whole cents. ## Disputes [Section titled “Disputes”](#disputes) If you dispute a payment with your bank instead, the disputed amount is removed from your balance and new work is paused with `402 PaymentDisputed` until the dispute closes. Running work continues. Contacting support first is usually faster. # Usage and costs > See what each run cost, split by project, label, meter and segment, and export usage as CSV. Every charge Nodus makes is backed by usage records: one line per meter, per object, per window. This page shows how to read them, group them and export them. ## See what a run cost [Section titled “See what a run cost”](#see-what-a-run-cost) Each compute object shows its cost so far in `status.cost`: ```bash nodus get job train-llama -o jsonpath='{.status.cost}' ``` ```json {"totalUSD": "4.182500", "heldUSD": "0.750000", "bySegment": {"bootUSD": "0.121000", "runningUSD": "3.980000", "teardownUSD": "0.081500"}} ``` `heldUSD` is reserved but not yet charged. The total becomes final once the capacity behind the run is confirmed deleted, which can be a minute or two after the run itself ends. When a run fails because of Nodus, the teardown after the failure and any of its start not yet billed are not charged: `bySegment.coveredByNodus` names those segments, and `nodus describe` shows them on the cost line, for example `$0.00 (covered by Nodus: boot, teardown)`. ## What each segment covers [Section titled “What each segment covers”](#what-each-segment-covers) Compute usage from the moment the capacity starts billing until it is confirmed deleted is split into segments: | Segment | Covers | | ---------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `Boot` | From the start of billing to the moment your command starts: boot, registration, readiness checks and pulling your image | | `Running` | Your command running, including reading inputs, lazy image fetches, checkpoint writes and the snapshot when the run stops | | `Restore` | For a run that resumes from a checkpoint or snapshot: from the start of billing to the moment your command starts again, including the restore. It replaces `Boot` | | `Teardown` | From the stop to confirmed deletion, plus the capacity’s billing increment (lines with `detail: IncrementRounding`) | Time between runs that reuse the same capacity is billed on the `warm_idle_seconds` meter and has no segment. When Nodus itself delays a deletion, that time is not charged: it shows as a zero-amount line with `detail: NodusBorne`. ## List and group usage [Section titled “List and group usage”](#list-and-group-usage) ```bash nodus get usage --since 2d nodus get usage --group-by project,meter nodus get usage --group-by member --since 30d nodus get usage --group-by label:team,day nodus get usage --group-by rank,segment --field-selector object.uid= ``` `--group-by` accepts `project`, `kind`, `member`, `meter`, `day`, `segment`, `rank` and `label:`. `member` is the person who started the work, whichever of their API keys they used. Labels are the ones the object carried when it was admitted, so label your work before you submit it: ```yaml metadata: name: train-llama labels: team: research ``` Grouped totals always add up to the ungrouped total: every line lands in exactly one group, and lines without the grouped value share an empty group. ## Export as CSV [Section titled “Export as CSV”](#export-as-csv) ```bash nodus get usage --since 30d -o csv > usage.csv curl -H "Authorization: Bearer $NODUS_TOKEN" -H "Accept: text/csv" \ "https://api.nodus-compute.ai/v1/usagerecords?since=30d" ``` Amounts are in US dollars with six decimals. `quantity` is in the meter’s unit (seconds, tokens, GiB or characters), and `rate_micros` is the rate in micro-dollars per `rate_basis` units. ## Inference and Indra [Section titled “Inference and Indra”](#inference-and-indra) Inference requests are rolled up into one line per model, meter and hour. A request to `nodus/indra` (Indra) shows the model that served it and, beside it, the routing call’s input and output tokens on lines whose SKU ends in `:routing:input` and `:routing:output`. The request is charged once, rounded up to the micro-dollar over all of its lines together. ## Agents [Section titled “Agents”](#agents) An Agent’s idle workers bill as warm idle on the Agent. When a run claims a worker, the worker’s time from that moment bills on the AgentRun until the run releases it, so each second of worker time appears once. Model calls a run makes with Nodus’s Claude are billed per request on the AgentRun, with the routing call beside the model that answered; calls made with your own Anthropic key carry no Nodus model charge. ## Storage and egress [Section titled “Storage and egress”](#storage-and-egress) Retained storage (checkpoints, Volumes, images you build, and Job and Agent outputs) is sampled every hour per object. The first 10 GB across your organization are included and shared across your objects in proportion to their size; each object’s line shows only its billable part. The hours of a UTC day are charged together shortly after midnight UTC, as one `Storage` transaction. Egress from your containers and interactive sessions is counted per object per UTC day. The first 10 GiB a day across your organization are included, and the rest is charged the next morning as one `Egress` transaction, split across objects by bytes. If your balance cannot cover a daily storage or egress charge, the remainder becomes arrears. While arrears are outstanding, new runs and uploads are refused with `ArrearsOutstanding`; your next top-up pays them first. Storage left in arrears gets an email at 7, 21 and 28 days. At 30 days Nodus proposes deleting the stored data beyond the included 10 GB: finished Jobs’ checkpoints and outputs first, then older Volume revisions, then the latest revisions, oldest first. Nothing is deleted until two Nodus administrators approve the list, and paying your arrears before then cancels it. Deleting data does not clear the arrears. ## Meters reference [Section titled “Meters reference”](#meters-reference) `quantity` is in the meter’s unit. `rate` applies to `rateBasis` units of quantity, so a line’s amount is `quantity × rate ÷ rateBasis`, rounded down, except that the last line of a run rounds up to the capacity’s billing increment. | Meter | Billed for | Unit | Rate is per | `rateBasis` | SKU | | ------------------------------ | ---------------------------------------------------------------------------------------------------------- | ------------------ | ------------------ | ----------- | -------------------------------------------------------------- | | `compute_seconds` | Dedicated capacity for Jobs, GPU Sandboxes, GPU Workspaces, Functions and training members, with a segment | seconds | hour | 3 600 | `rented:` | | `warm_idle_seconds` | Warm capacity kept between runs, idle Function and agent workers | seconds | hour | 3 600 | `rented:` | | `node_vcpu_seconds` | vCPU on shared nodes (Sandboxes, agent runs), with a segment | milli-vCPU seconds | vCPU-hour | 3 600 000 | `node:vcpu` | | `node_gib_seconds` | Memory on shared nodes, with a segment | MiB seconds | GiB-hour | 3 686 400 | `node:memory` | | `node_disk_gib_seconds` | Requested disk above 10 GiB per vCPU on shared nodes | GiB seconds | GiB-hour | 3 600 | `node:disk` | | `build_vcpu_seconds` | Image builds, with a segment | milli-vCPU seconds | vCPU-hour | 3 600 000 | `build:vcpu` | | `storage_gb_hours` | Retained storage above 10 GB per organization | MB hours | GB-month (30 days) | 720 000 | `storage:retained` | | `egress_gib` | Egress above 10 GiB per organization per UTC day; relayed training traffic from the first byte | MiB | GiB | 1 024 | `egress:gib`, `egress:mesh-relay`, `egress:supplier` | | `inference_input_tokens` | Input tokens | tokens | million tokens | 1 000 000 | `model::input` | | `inference_output_tokens` | Output tokens, reasoning included | tokens | million tokens | 1 000 000 | `model::output` | | `inference_cache_read_tokens` | Cached input read | tokens | million tokens | 1 000 000 | `model::cache_read` | | `inference_cache_write_tokens` | Cache writes, 5-minute or 1-hour | tokens | million tokens | 1 000 000 | `model::cache_write_5m`, `model::cache_write_1h` | | `inference_audio_seconds` | Transcription and translation, 10 s minimum per request | milliseconds | audio-hour | 3 600 000 | `model::audio` | | `inference_speech_characters` | Speech synthesis input | characters | million characters | 1 000 000 | `model::speech` | | `platform_device_hours` | Your own devices assigned to Nodus-scheduled work | device seconds | device-hour | 3 600 | `platform:device-hour` | | `platform_predict_seconds` | Predict on one of your pools | seconds | 30-day month | 2 592 000 | `platform:predict` | Indra requests add a routing line beside the model’s lines, with SKU `:routing:input` or `:routing:output`. An `egress:supplier` line passes through an egress charge billed for your dedicated capacity at the same margin as the capacity itself. # What you pay for > Every kind of time and cost, who pays for it, and the usage segment it appears under. You pay for every second a provider bills for a machine Nodus acquired for your work, from the moment billing starts until the machine is confirmed deleted. Nodus pays for everything it chose or caused. Your usage shows each machine’s time by segment, so you can see where every second went: * **Boot**: from billing start to your command starting. * **Restore**: a replacement machine’s time until your command starts again after a recovery, including the checkpoint restore. * **Running**: from your command starting to its stop. * **Teardown**: from stop to confirmed deletion, plus the provider’s billing increment, rounded once per machine. Run `nodus get usage --group-by segment` or open **Usage & billing → Usage** in the console to see the split. For a multi-node run, `--group-by rank,segment` itemizes each member. ## The table [Section titled “The table”](#the-table) | What | Segment | Who pays | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------- | -------- | | []()Time from billing start to your command starting: boot, registration, readiness checks, image pull and checkpoint restore | Boot, Restore | You | | []()A boot that fails on the provider’s side, until the machine is confirmed deleted | Boot | You | | []()Spare machines Nodus starts to finish your work sooner, machines Nodus releases or replaces on its own decision, machines that fail Nodus’s checks, and failures Nodus causes | None: never on your usage | Nodus | | []()From your command starting to its stop, including lazy image fetches, checkpoint writes, the snapshot on stop and work lost after a preemption | Running | You | | []()Downloading inputs after your command starts, and the `initCommand` preflight | Running | You | | []()Imports, exports and sink loads that run on your org’s capacity | Running | You | | []()Disk above 10 GiB per vCPU on Nodus nodes | Running | You | | []()A network partition your work tolerates, until it reconnects or until the old machine is fenced and confirmed deleted | Running, Teardown | You | | []()The provider’s billing increment, rounded once per machine | Teardown | You | | []()Warm idle time beyond an increment you already paid for, minimum workers, and idle Function and agent workers | Warm idle | You | | []()The idle share of Nodus-operated capacity, built into its rates and never metered on its own | Included in rates | You | | []()From stop to confirmed deletion | Teardown | You | | []()Time between stop and the delete request beyond 60 seconds when Nodus caused the delay | None: never on your usage | Nodus | | []()Any cost beyond your tightest limit: `maxCostUSD`, a Budget or your balance | None: never on your usage | Nodus | | []()A model request whose outcome Nodus cannot confirm, and upstream charges beyond the usage observed on a cut-off stream | None: never on your usage | Nodus | | []()Console assistant and command generation calls, within the daily cap | None | Nodus | | []()Grading Sandboxes, evaluations, dataset previews, task manifests and model calls made from inside your containers, under your caps | Running | You | | []()Mirroring a private image so it can run on hosted containers, and the mirror’s storage | None | Nodus | | []()Nodus’s own test runs | None | Nodus | | []()Image builds that fail | Build | You | | []()Storage above 10 GB per org: Volumes, checkpoints, outputs, images and session state | Storage | You | | []()Logs | None | Nodus | | []()Egress within 10 GiB per org per day | None | Nodus | | []()Egress beyond the included 10 GiB a day, up to your quota, and relayed traffic between multi-node members | Egress | You | | []()Your own pool hosts and cloud accounts | None | Free | | []()Pool devices running work Nodus schedules, Predict per pool, and capacity Nodus acquires for you | Platform | You | | []()Differences between a provider’s invoice and what Nodus metered | None | Nodus | This table is checked against the billing reference model on every change, so the rows here are the rows Nodus charges by. ## Limits stop work before it overruns [Section titled “Limits stop work before it overruns”](#limits-stop-work-before-it-overruns) Every paid action holds funds before it starts, and work stops gracefully when a hold cannot be renewed. If a charge would still go past your tightest limit, Nodus absorbs the difference: your balance never pays more than the limit you set. See [How billing works](/docs/concepts/billing/) for holds, captures and the low-balance stop. # When money runs out > How Nodus warns you, stops work gracefully and resumes it when credits, a raised Budget or a raised cap return. Credits are prepaid. Before paid work starts, Nodus places a funded hold on your balance, and it renews that hold every 5 minutes while the work runs. When a renewal cannot be funded, from your balance, a Budget or a `maxCostUSD` cap, the work stops gracefully inside time that is already paid for. Nothing runs past what you funded, and nothing is lost that the work saved. ## Before you start: fail fast [Section titled “Before you start: fail fast”](#before-you-start-fail-fast) A create that cannot be funded is refused at once with the amounts and the fix: ```text 402 InsufficientCredits job "train-a" needs a $3.20 hold to start (released when it ends); available $1.10. fix: nodus billing top-up 20, or lower spec.maxCostUSD ``` Other refusals are `BudgetExceeded` (a Budget or a cap has too little left), `ArrearsOutstanding` (unpaid charges) and `PaymentDisputed`. A distributed Job checks all of its nodes together: it starts with every node funded or not at all. ## Warnings [Section titled “Warnings”](#warnings) When your available balance falls below $5, or below 20 % of what the next renewal of your running work needs, you get a `LowBalance` Event, a `billing.low_balance` webhook, at most one email a day and a console banner. Running work whose funds will run out soon still shows `Funded=True`, and the condition’s message gives the time it will stop. A top-up, a credit grant or an auto-recharge before that time keeps it running. When a charge is larger than what is left, the rest becomes unpaid charges. You get an `ArrearsPosted` Event, a `billing.arrears_posted` webhook and at most one email a day, and new work waits until a top-up or credit grant pays them. ## How each kind stops [Section titled “How each kind stops”](#how-each-kind-stops) | Kind | What happens | State | Resumes | | ------------------------------ | ----------------------------------------------------------------------------------- | -------------------------------------------------------------------------- | --------------------------------------------------------- | | Job (checkpointed) | Takes an urgent checkpoint, then stops | `Suspended`, `Funded=False` | Automatically once funded, from the checkpoint | | Job (restartable or ephemeral) | Stops at the funded edge | `Suspended`, `Funded=False` | Automatically; ephemeral Jobs restart from the start | | Job at `maxCostUSD` | Checkpoints, then stops | `Suspended`, reason `MaxCostReached` | When you raise `maxCostUSD` | | Distributed Job | Every node checkpoints together, then all stop | `Suspended`, `Funded=False` | Automatically, with all nodes funded | | Sandbox | Snapshots its filesystem and state, then stops | `Stopped`, `stopReason: InsufficientCredits`; `spec.state` stays `Running` | On the next exec or connect once funded | | Workspace | Saves the home volume, then stops | `Stopped` with the money reason | On `nodus start`, SSH connect or the schedule once funded | | AgentRun | Finishes or interrupts the current step | `Waiting` for funds | Automatically | | Agent idle workers | Scale to zero | `Funded=False` on the Agent | Automatically | | Function | Stops taking calls, finishes in-flight calls, scales to zero | New calls queue with `Funded=False` | Automatically | | TrainingJob | Its Jobs suspend; grading Sandboxes are deleted | `Suspended` | Automatically | | Image build | The build stops | Image `Pending`, `Funded=False` | Automatically rebuilds | | Volume import or export | The transfer stops | `ImportFailed(InsufficientCredits)` | `nodus request reimport volume/` | | Inference | New requests return an OpenAI- or Anthropic-shaped `402`; in-flight requests finish | — | Immediately | Nodus never changes the `spec.state` you set. It records the stop in the object’s phase and `Funded` condition, emits `FundingLost`, and emits `FundingRestored` when it resumes. You get at most one email an hour listing stopped work. ## Getting going again [Section titled “Getting going again”](#getting-going-again) ```bash nodus billing top-up 20 nodus get jobs --watch ``` A top-up resumes waiting work oldest first, as each hold fits. Work stopped by a Budget resumes when you raise the Budget’s limit, switch it to `Notify`, delete it or the next month starts; work stopped at `maxCostUSD` resumes when you raise the cap. # Checkpoints > Keep a Job's progress through suspends, stops and lost capacity by saving it to the state directory, and answer checkpoint requests so every save is consistent. A checkpoint is a copy of your Job’s **state directory** that Nodus takes while the Job runs. When the Job is suspended, or the capacity under it is reclaimed, the next attempt starts with that directory restored and your program carries on from what it saved. Nodus decides when to checkpoint and where to store it; your program decides what to write and how to load it. Checkpoints hold files, not memory. A restored attempt starts your command again from the beginning, with the state directory as it was at the last checkpoint, so the program must read its own progress back. ## Save your progress to the state directory [Section titled “Save your progress to the state directory”](#save-your-progress-to-the-state-directory) Write everything you need to continue, such as model weights, optimizer state and the current step, under `/nodus/state`. The path is also in `NODUS_STATE_DIR` (`NODUS_CHECKPOINT_DIR` is the older name for the same directory). On start, load what is there: examples/checkpoints/resume/train.py ```python """A training loop that saves its progress where Nodus checkpoints it and resumes from it. It uses only the standard library, so it runs in any image. With the Nodus SDK installed, `nodus.state_dir()`, `nodus.checkpoint.on_request()` and `nodus.restored()` do the same work. """ import json import os import socket import threading import time from pathlib import Path STATE_DIR = Path(os.environ.get("NODUS_STATE_DIR", "/nodus/state")) STATE = STATE_DIR / "progress.json" STEPS = int(os.environ.get("STEPS", "120")) SAVE_EVERY = 20 # like Hugging Face Trainer's save_steps: a regular save even when Nodus does not ask pending = threading.Event() # set while Nodus waits for a consistent checkpoint request_seq = None events = None def save(step): """Write the state atomically, so a snapshot never sees a half-written file.""" STATE_DIR.mkdir(parents=True, exist_ok=True) tmp = STATE.with_suffix(".tmp") tmp.write_text(json.dumps({"step": step})) os.replace(tmp, STATE) def listen(): """Subscribe to checkpoint requests on the events socket and flag each one for the loop.""" global events, request_seq path = os.environ.get("NODUS_EVENTS_SOCKET", "/run/nodus/events.sock") try: events = socket.socket(socket.AF_UNIX, socket.SOCK_STREAM) events.connect(path) except OSError: return # outside Nodus there is nobody to ask for checkpoints events.sendall(b'{"type":"checkpoint.subscribe"}\n') for line in events.makefile("r"): message = json.loads(line) if message.get("type") == "checkpoint.request": request_seq = message["seq"] pending.set() def ack(): """Tell Nodus the files in the state directory are complete.""" events.sendall((json.dumps({"type": "checkpoint.ack", "seq": request_seq}) + "\n").encode()) pending.clear() def main(): start = 0 if STATE.exists(): start = json.loads(STATE.read_text())["step"] print(f"resumed from step {start}", flush=True) threading.Thread(target=listen, daemon=True).start() for step in range(start + 1, STEPS + 1): time.sleep(1) # one step of work print(f"step {step}/{STEPS}", flush=True) if pending.is_set(): save(step) ack() elif step % SAVE_EVERY == 0: save(step) print("training complete", flush=True) if __name__ == "__main__": main() ``` Two habits keep checkpoints useful: * **Write atomically.** Write to a temporary file and rename it over the old one, as `save()` does. A checkpoint can then never capture a half-written file. * **Keep outputs separate.** Results you want to download go to `/nodus/outputs`. The state directory is recovery state: it is restored into the next attempt, not offered as a download. `NODUS_RESTORED=1` is set in an attempt that started from a checkpoint, if your program wants to log the difference. An empty state directory never replaces an earlier checkpoint that had files in it, so an attempt that fails before it saves anything cannot erase progress. ## Choose what is saved [Section titled “Choose what is saved”](#choose-what-is-saved) `recovery.checkpoint` in the Job spec controls the checkpoint. The defaults suit most programs: ```yaml recovery: continuity: Checkpointed # restore the latest checkpoint into each new attempt checkpoint: paths: [/nodus/state] # up to 64 absolute paths, all saved in one checkpoint interval: auto # or a fixed interval from 1m to 6h maxSize: 1Ti # larger checkpoints fail with CheckpointTooLarge retainAfterFinish: 168h # after the Job finishes, keep only the final checkpoint ``` On the command line, `nodus run --checkpoint /nodus/state` sets `paths`. Saving a whole folder such as your working directory is possible by listing it, but it makes every checkpoint larger and slower; list only what you need to continue. `continuity: Restartable` is the lighter choice for programs that track their position as a counter: Nodus keeps `NODUS_CURSOR_COMPLETED` and `NODUS_CURSOR_TOTAL` from your progress reports and skips the file copy. `continuity: Ephemeral` starts every attempt from scratch. ## When Nodus checkpoints [Section titled “When Nodus checkpoints”](#when-nodus-checkpoints) With `interval: auto`, Nodus sets the cadence from how often the capacity your Job runs on is interrupted and how long a checkpoint takes to save. A Job gets at least four checkpoints over its expected run time, and checkpointing takes no more than about a tenth of it. Capacity that is rarely interrupted is checkpointed less often. Nodus also checkpoints, whatever the interval: * when you run `nodus suspend job/NAME`, before the compute is released, so `nodus resume` continues from there; * when the capacity gives notice that it is about to be reclaimed, so the replacement attempt loses as little work as possible. `nodus describe job/NAME` shows the latest checkpoint, and `GET …/jobs/NAME/checkpoints` lists each one with its sequence number, attempt, time, size and file count. ## Answer checkpoint requests [Section titled “Answer checkpoint requests”](#answer-checkpoint-requests) A checkpoint taken while your program is halfway through writing its files would restore a broken state. To avoid that, a program can ask to be told before each checkpoint and say when its files are complete. This is the **request/ack handshake**, and the example above implements it with the standard library. It runs over the events socket at `/run/nodus/events.sock` (`NODUS_EVENTS_SOCKET`), one JSON object per line: 1. Your program connects and sends `{"type": "checkpoint.subscribe"}` once. 2. Before each checkpoint, Nodus sends `{"type": "checkpoint.request", "seq": 7, "urgent": false}`. 3. Your program finishes the current step, writes its state and replies `{"type": "checkpoint.ack", "seq": 7}`. 4. Nodus copies the state directory, then your program carries on. It does not need to pause while the copy runs. `urgent: true` means the capacity is about to go away: save at the next safe point and skip optional work. Nodus waits for the ack for as long as the shutdown allows and then takes the checkpoint anyway, so a program that hangs cannot block a suspend. You choose whether to use the handshake with `recovery.checkpoint.integration`: | Value | Behaviour | | ---------------- | --------------------------------------------------------------------------------------------------------------- | | `Auto` (default) | Use the handshake when the program subscribes; otherwise checkpoint without asking | | `None` | Never ask; checkpoint the paths as they are | | `HFTrainer` | Answer requests from inside Hugging Face Trainer, with no change to your image ([below](#hugging-face-trainer)) | With the Python SDK installed (`pip install nodus-compute`), `nodus.checkpoint.on_request(save)` registers a callback and sends the ack after it returns, `nodus.checkpoint.requested()` lets a loop poll instead, and `nodus.state_dir()` returns the directory. All of them do nothing outside Nodus, so the same script runs on your laptop. ## Hugging Face Trainer [Section titled “Hugging Face Trainer”](#hugging-face-trainer) Trainer already saves and resumes; point it at the state directory and it works with Nodus checkpoints: * set `output_dir` to the state directory (`os.environ["NODUS_STATE_DIR"]`), or a folder inside it; * set `save_steps` to how often Trainer saves on its own, and `save_total_limit` (for example `2`) so old checkpoints do not fill the directory; * call `trainer.train(resume_from_checkpoint=True)` when `NODUS_RESTORED` is `1`, and `trainer.train()` on a first start, because Trainer refuses to resume from an empty directory. To also save when Nodus asks, add `nodus.checkpoint.HFTrainerCallback()` to the Trainer’s callbacks, or set `integration: HFTrainer` and Nodus registers the same callback for you, even in an image without the SDK. Trainer then saves at the end of the current step and Nodus checkpoints once the save is written. ## Gang checkpoints (Beta) [Section titled “Gang checkpoints (Beta)”](#gang-checkpoints-beta) Beta Multi-node Jobs (`distributed`) are Beta. Their checkpoints use a different format. A Job with `distributed` set defaults to `recovery.checkpoint.format: Dcp`. Every rank writes its shard with `torch.distributed.checkpoint` to `$NODUS_CHECKPOINT_URI`, and rank 0 commits it, which `nodus.checkpoint.dcp.save` and `.load` do for you. After a restart, `$NODUS_RESTORE_URI` names the latest committed checkpoint. A rank that lost its place in the gang cannot commit, so a restored gang always loads a checkpoint every rank finished. ## Run the example [Section titled “Run the example”](#run-the-example) The example suspends the Job mid-run and resumes it; the second attempt prints `resumed from step N` and finishes the remaining steps: ```console $ cd examples/checkpoints/resume $ nodus run --name checkpoints-resume --cpu 2 --checkpoint /nodus/state -d -- python train.py $ nodus suspend job/checkpoints-resume $ nodus resume job/checkpoints-resume $ nodus logs -f job/checkpoints-resume resumed from step 12 step 13/120 … ``` Checkpoints are deleted with their Job. After a Job finishes, Nodus keeps only its final checkpoint once `retainAfterFinish` (seven days by default) has passed. # Which Nodus compute feature should I use? > Compare Jobs, Workspaces, Sandboxes and Functions. Choose the right way to run code, call models, run agents or train a model on Nodus. Use a **Job** when you have a command that should run to completion. Use a **Workspace** for interactive GPU or CPU development, a **Sandbox** for isolated CPU code execution, and a **Function** to call Python remotely or in parallel. Nodus also provides hosted inference, managed agents and training workflows. ## Choose by the work you need to do [Section titled “Choose by the work you need to do”](#choose-by-the-work-you-need-to-do) | I need to… | Start with | Why it fits | | ----------------------------------------------------- | ------------------------------------------- | --------------------------------------------------------------------------------- | | Run a script, batch task or existing training command | [Jobs](/docs/guides/jobs/) | Runs a command to completion with logs, outputs and resource and cost limits | | Develop with SSH, VS Code or JupyterLab | [Workspaces](/docs/guides/workspaces/) | An interactive GPU or CPU machine with a saved home directory | | Execute agent-generated or untrusted code | [Sandboxes](/docs/guides/sandboxes/) | An isolated CPU container driven by command and file requests, with idle stopping | | Call Python remotely or map it over many inputs | [Functions](/docs/guides/functions/) | Remote calls with workers that scale with demand | | Call a hosted language model | [Inference (Beta)](/docs/guides/inference/) | An OpenAI-compatible model API billed per token | | Run an agent conversation with shell tools | [Agents (Beta)](/docs/guides/agents/) | An AgentRun with its own sandbox and recorded execution history | | Fine-tune or train through managed runtimes | [Training (Beta)](/docs/guides/training/) | A TrainingJob defines the model, data, runtime and training parameters | For a first run, follow [your first job](/docs/getting-started/). It walks through signing in, running a command, following logs and downloading results. ## What is the difference between a Job and a Function? [Section titled “What is the difference between a Job and a Function?”](#what-is-the-difference-between-a-job-and-a-function) A Job runs an executable command and finishes when that command exits. A Function is a Python callable deployed in an App; your code invokes it with `.remote()`, `.map()` or `.spawn()`. Choose Jobs for an existing script or batch process. Choose Functions when remote calls should be part of your Python application. Both can use GPU or CPU workers. Function workers can stay warm between calls; read [Function billing](/docs/guides/functions/billing/) before choosing idle and scaling settings. ## What is the difference between a Workspace and a Sandbox? [Section titled “What is the difference between a Workspace and a Sandbox?”](#what-is-the-difference-between-a-workspace-and-a-sandbox) A Workspace is for interactive development with SSH, VS Code or JupyterLab on GPU or CPU compute. Its home directory is backed by a Volume. A Sandbox is an isolated CPU container for programmatic commands and file operations, including code produced by an agent. Its network is closed unless you open it, and it can stop after an idle period. Read [Workspace storage](/docs/guides/workspaces/) and [Sandbox isolation](/docs/concepts/sandboxes-isolation/) before deciding what state and access your task needs. ## Do I need Agents to use Claude Code, Codex or Cursor? [Section titled “Do I need Agents to use Claude Code, Codex or Cursor?”](#do-i-need-agents-to-use-claude-code-codex-or-cursor) No. Connect your existing coding agent to Nodus through [MCP](/docs/guides/mcp/) using the [client setup guide](/connect/). It can work with Nodus resources through that connection. The Agents feature is for running an agent conversation inside Nodus itself. ## What survives an interruption? [Section titled “What survives an interruption?”](#what-survives-an-interruption) Recovery depends on the feature and its configuration. For checkpointed Jobs, your program writes and loads its own files in `NODUS_CHECKPOINT_DIR`; Nodus saves and restores that directory. This does not restore arbitrary process or GPU memory. Restartable Jobs start over, while ephemeral Jobs keep no state. Read [Checkpoints](/docs/guides/checkpoints/) for application recovery, [Volumes](/docs/guides/volumes/) for persistent files and [Outputs](/docs/guides/outputs/) for results you need to download. ## How do I control spending? [Section titled “How do I control spending?”](#how-do-i-control-spending) Check the estimate before starting work, set the resource’s supported cost and time limits, and configure [project budgets](/docs/guides/billing/budgets/). Compute usage, saved storage and model tokens have different billing rules; the [billing guide](/docs/guides/billing/) explains them. Use the [pricing reference](/docs/reference/pricing/) for published amounts instead of copying prices from an old example. ## Where should an agent look up exact syntax? [Section titled “Where should an agent look up exact syntax?”](#where-should-an-agent-look-up-exact-syntax) Use the [CLI reference](/docs/reference/cli/) for commands, [Python SDK reference](/docs/reference/python/) for signatures and [OpenAPI](/docs/openapi.json) for v1 request and response fields. Beta resources use the [v1beta1 contract](/docs/openapi-v1beta1.json). Every authored guide has a Markdown version with its examples. The [agent guide](/docs/for-agents/) links a lightweight page manifest and focused topic bundles for retrieval. # Connections > Connect databases, S3 buckets, Weights & Biases and GitHub once, verified, and use them from Jobs, Sandboxes, imports and outputs. A Connection is an external system that Nodus talks to for you: a Postgres database for query imports and outputs, an S3 bucket for inputs, a Weights & Biases project for live run tracking, or your GitHub repositories for private sources. Its credentials live in a [Secret](/docs/guides/secrets/); the Connection says what they are for, and Nodus checks that they work before anything uses them. ## Create a Connection [Section titled “Create a Connection”](#create-a-connection) Put the credentials in a Secret, then point the Connection at it: ```console nodus secret create analytics-db --from-literal DATABASE_URL=postgres://reader:...@db.example.com:5432/app ``` connection.yaml ```yaml apiVersion: nodus.dev/v1 kind: Connection metadata: name: ex-connections-postgres spec: type: Postgres secret: ex-connections-postgres # holds DATABASE_URL scope: Read # verified: the role can read, and is not asked to write ``` ```console nodus apply -f connection.yaml nodus wait connection/ex-connections-postgres --for=condition=Verified nodus describe connection/ex-connections-postgres ``` | `type` | Secret keys | Settings | Verified by | | ------------------ | -------------------------------------------------------------- | ---------------------------------------------------- | -------------------------------------------------------------------- | | `Postgres` | `DATABASE_URL` | `scope`: `Read`, `Write` or `ReadWrite` | Connecting, a query, and for write scopes the right to create tables | | `Neon`, `Supabase` | `DATABASE_URL` | `scope`; `neon.branch` | The same checks, on the provider’s own host | | `S3` | `ROLE_ARN`, or `AWS_ACCESS_KEY_ID` and `AWS_SECRET_ACCESS_KEY` | `s3.bucket`, `s3.region`, `s3.prefix`, `s3.endpoint` | Reaching the bucket with short-lived credentials | | `WandB` | `WANDB_API_KEY` | `wandb.entity`, `wandb.project`, `wandb.live` | The key’s access to the entity | Prefer an IAM role (`ROLE_ARN`) for S3: Nodus assumes it for minutes at a time, so no long-lived key is stored. Every S3 Connection has its own external id in `status.externalId`, and Nodus sends exactly that id each time it assumes the role, so a role your trust policy grants to one Connection can’t be used through any other Connection or organization. Create the Connection, read the id, and require it in the role’s trust policy: ```console nodus get connection/datasets -o jsonpath='{.status.externalId}' ``` ```json { "Effect": "Allow", "Principal": { "AWS": "" }, "Action": "sts:AssumeRole", "Condition": { "StringEquals": { "sts:ExternalId": "" } } } ``` The Connection shows `Failed` until the trust policy names the id; the next check, within the hour, turns it `Ready`. You don’t need `EXTERNAL_ID` in the Secret. If you set it, it must equal `status.externalId`, and any other value is refused. A Connection is `Ready` once verified, and `status.egressHosts` lists the hosts it may reach. Nodus checks it again every day, and every hour after a failure. When a check fails the Connection shows `Failed`, the `Verified` condition says why, and anything that needs it fails with `ConnectionNotReady` until you fix the Secret. ## Use it [Section titled “Use it”](#use-it) ```yaml spec: connections: [tracking] # injects WANDB_* and allows its hosts through the egress policy inputs: - name: shards bucket: {uri: s3://acme-data/shards/, connection: datasets} ``` A `WandB` Connection with `wandb.live: true` sets `WANDB_API_KEY`, `WANDB_ENTITY`, `WANDB_PROJECT`, `WANDB_RUN_GROUP`, `WANDB_NAME` and `WANDB_RUN_ID` for the Job, so `wandb.init()` needs no arguments. Each index of the Job logs to one run, and a run that Nodus restarts after lost capacity resumes it rather than starting a new one. A Job with one index shows the run’s page in `status.links`. A variable you set yourself keeps your value; setting `WANDB_RUN_ID`, `WANDB_ENTITY` or `WANDB_PROJECT` yourself means Nodus doesn’t link the run. A Job that names a Connection that is missing or not `Ready` runs without it and gets a `ConnectionNotReady` Event. ## Connect GitHub [Section titled “Connect GitHub”](#connect-github) A `GitHub` Connection installs the Nodus GitHub App on your account or organization. It needs no Secret: ```console nodus create connection github --type GitHub nodus get conn github -o jsonpath='{.status.installURL}' ``` Open the link, choose the repositories to share and install the App. GitHub returns you to Nodus, which checks that you can see the installation, and the Connection becomes `Ready` with `status.github` naming the account and repositories. `GET …/connections/github/repositories?limit=100` pages through all of them. A Sandbox’s `init.git` can then clone a private repository: Nodus resolves the branch or tag to a commit when you create the Sandbox and clones that commit. The clone uses a short-lived token that can only read that one repository, and the token never appears in the Sandbox’s environment or files. If someone uninstalls the App, the Connection shows `Failed` with `InstallationRemoved` and a fresh `status.installURL` to install it again. A Volume can import from a bucket (`source.s3`) or a query (`source.connectionQuery`); see [Volumes](/docs/guides/volumes/#import-from-elsewhere). A Job or Sandbox reaches only the hosts its Connections were verified against. # Use the console > Sign in, find your work, follow it live and copy the matching CLI command from any page. The console at `console.nodus-compute.ai` shows the same objects the CLI and the Python SDK work with. Every page has the command that does the same thing, so you can move between them at any time. ## Sign in [Section titled “Sign in”](#sign-in) 1. Open the console and sign in with your email and password, or with GitHub or Google. 2. New accounts confirm their email with a 6-digit code. 3. If someone invited you, the console shows the invite first. Accept it to join their organization. 4. Otherwise create your organization. Its URL name (for example `research`) appears in every link. Your first organization receives starter credit once your email is verified. ## Sign in the CLI from the browser [Section titled “Sign in the CLI from the browser”](#sign-in-the-cli-from-the-browser) Run `nodus login`. The CLI opens the console, which asks which organizations the CLI may use and names the key it creates, for example `cli-laptop-2026-09-30`. Choose **Authorize** and return to your terminal. On a machine without a browser, run `nodus login --device` and enter the code it prints on the console’s device page. If you have no organization yet, authorizing creates one for you. If someone invited you, both pages show the invite first: accept it, and the CLI signs in to their organization. ## Find your work [Section titled “Find your work”](#find-your-work) * The URL always names the organization and project: `/research/default/jobs/lora-fine-tune/logs`. Links you share open in the same place for anyone with access, and two browser tabs can work in two organizations. * Switch organization or project from the top bar. * Press `Cmd+K` (or `Ctrl+K`) to search objects, pages and actions. Type `job/` or `sb/` and a name to jump straight to an object. * Press `?` for every keyboard shortcut. `g` then a letter goes to a section, for example `g b` for Billing. ## Follow runs live [Section titled “Follow runs live”](#follow-runs-live) Lists and detail pages update as things change, without reloading. The note **Live · 12s ago** shows the time since the last update. Choose it and **Pause live updates** to freeze what you are reading; **Resume** applies the changes at once. Running work shows its cost so far, marked with `≈` between updates. ## Copy the command [Section titled “Copy the command”](#copy-the-command) The terminal icon in each page header copies the command for what you are looking at, with the project and organization spelled out: ```sh nodus describe job/lora-fine-tune -p default --org research ``` **View YAML** in the actions menu shows the live object as a manifest you can download and apply elsewhere. ## Create something [Section titled “Create something”](#create-something) Create pages show four tabs: **Form**, **YAML**, **CLI** and **Python**. They always show the same manifest. Before you launch, the estimate shows the expected cost range, the hourly rate, the first hold and how long the price holds. Launch uses exactly that estimate; if prices move first, the console estimates again and asks you to confirm. Your form is saved as a draft in this browser until you launch or discard it. If the network drops while you launch, the console says it could not confirm the result. **Check status** looks for the object by name, and **Retry** is safe: it can never create a second copy. ## When credits run out [Section titled “When credits run out”](#when-credits-run-out) If your balance cannot fund the next renewal, running work stops at its funded edge. The project Overview then shows **Paused for funds**, grouped by what adding credits does: * **Resumes automatically**: jobs and agent runs continue by themselves once a hold fits. * **Needs a start**: workspaces stay stopped until you choose **Start** (or **Start all**). * **Wakes on next use**: sandboxes start again on the next command, file request or preview. * **Rejecting requests**: inference endpoints answer with `402` until funded. * **Failed imports**: choose **Retry import** on the volume. * **Needs a higher limit**: work that reached its own spend limit resumes only when you raise that limit. Choose **Add credits** on the panel to top up. ## Appearance [Section titled “Appearance”](#appearance) Choose **System**, **Light** or **Dark**, and **Comfortable** or **Compact** density, from the account menu. The theme follows you across the console and the docs. # Environments > Task sets with graders for reinforcement learning and evaluation, from the Nodus catalog or your own image. An Environment is a versioned set of tasks with a grader: a prompt for each task, and a program that decides whether a completion is correct. TrainingJobs use Environments for reinforcement learning and evaluation, and agent evaluations use the same ones. Environments are generally available (`nodus.dev/v1`). ## Browse the catalog [Section titled “Browse the catalog”](#browse-the-catalog) ```console $ nodus get environments -n nodus NAME VERSION CATEGORY MODES PHASE graph-coloring 1.0.0 Reasoning Train, Evaluate Ready arithmetic 2.0.0 Math Train, Evaluate Ready gsm8k 1.0.0 Math Train, Evaluate Ready reasoning-gym 1.0.0 Reasoning Train, Evaluate Ready python-functions 1.0.0 Code Train, Evaluate Ready $ nodus describe environment/graph-coloring -n nodus ``` `describe` shows the summary, licenses, the size of each split, a sample task, the graders and the examples: ready TrainingJob templates with measured results. The console’s Training page lists the same catalog, and “Run example” starts a TrainingJob from one; in Python, `nodus.TrainingJob.from_example("nodus/graph-coloring@1.0.0")` does the same. ## Use one in a TrainingJob [Section titled “Use one in a TrainingJob”](#use-one-in-a-trainingjob) Name the environment and version, and how many tasks to train and evaluate on: ```yaml spec: runtime: nodus/grpo-lora environment: name: nodus/graph-coloring@1.0.0 trainTasks: 50 # from the train split heldOutTasks: 64 # from the test split, never trained on seed: 42 ``` The same name, version and seed always give the same tasks in the same order, so two runs are comparable. The train and test splits never share a task. ## How grading works [Section titled “How grading works”](#how-grading-works) * Tasks carry a prompt and non-secret metadata only. Expected answers stay with the grader; the trainer, the model and your code never see them. * Completions are graded by Nodus, not by the trainer. Program graders run in grading Sandboxes that Nodus creates for your TrainingJob or agent evaluation: no network access, deleted after five minutes idle, at most `grading.maxParallel` at once. Code from a completion runs as a separate user that cannot read the answers. `ExactMatch` graders compare the completion with the answer inside Nodus, without a Sandbox. * Every verdict is `Correct`, `Incorrect`, `InvalidOutput` or `InfrastructureFailure`, with a reward and evidence recorded by the host that ran the grader: the process, its exit code, its duration and a hash of its output. * `InvalidOutput` is a completion the grader cannot parse, such as an answer without the expected tag. It is a failed task with a reward of 0. * `InfrastructureFailure` means the grader could not run. It carries no reward and does not count toward a pass rate, so an outage never looks like a wrong answer. * Grading Sandboxes are billed to the TrainingJob and count toward its `maxCostUSD`. They are deleted when the run finishes or is suspended. ## Train on a reward function in one call [Section titled “Train on a reward function in one call”](#train-on-a-reward-function-in-one-call) When your tasks and reward are in Python, `rl.train` is all you need. To see it work first, run the example that ships with the SDK: ```sh python -m nodus.examples.rl ``` It is one file, `nodus/examples/rl.py`, and it is the template for your own task: ```python import re from nodus.recipes import rl ANSWER = re.compile(r"\s*([A-Za-z]+)\s*") tasks = [(f"Spell the word backwards inside .\n\nWord: {w}", w[::-1]) for w in WORDS] def reward(completion, answer): found = ANSWER.findall(completion) return None if not found else float(found[-1].lower() == answer) run = rl.train(tasks, reward, max_cost=2) run.watch() # each stage, then the reward, loss and KL of every step run.outputs.download("./outputs") # the LoRA adapter and the before-and-after comparison ``` `run.tasks(phase="Evaluation", outcome="Failed")` lists the graded tasks the run has scored so far, each with its `taskId`, `reward` and outcome, so you can read which tasks the trained model still gets wrong. * **Tasks** are `(prompt, answer)` pairs or `{"prompt", "answer", "metadata"}` dicts. Only `reward` sees the answer. * **Model.** The base model is `Qwen/Qwen3-0.6B` unless you pass `model=`, for example `model="Qwen/Qwen3-4B"`. * **Held-out tasks.** A fifth of the tasks, and at least 16, is held out. The model is graded on them before and after training. Pass `test_tasks=` to choose them yourself. * **What gets packaged.** The reward goes with the lines of its own file that it uses: imports, constants and helper functions. A name it cannot take along fails before anything is built. * **Checked before it runs.** `rl.train` grades one task’s answer and an empty reply on your machine first. A reward that raises, or returns something other than a number, a bool or `None`, fails there instead of on a GPU. Return `None` or `0` when a reply holds no answer. * **Reuse.** The same code and tasks reuse the same Environment, so a second run starts at once. * **No network.** The reward runs in Nodus’s grader, which has no network access. Pass `pip=["package==1.2.3"]` for packages it imports. * **How it learns.** Each step samples `group_size=8` replies to each of `groups_per_step=4` tasks and scores every reply against the others in its group. Replies are capped at `max_tokens=256`, and the learning rate is `learning_rate=4e-5`. You can change any of them. * **Format rewards.** To reward the shape of an answer as well as its value, fold both into the reward, for example `return float(correct) - 0.1 * (not well_formed)`. * **Other options.** `steps`, `gpu`, `max_cost`, `lora` and the other `grpo_lora` parameters are optional. The run appears under **Runs › Post-training** in the console, with its curves, results and cost. To train on a catalog Environment, pass its name instead of tasks and a reward: ```python run = rl.train("nodus/gsm8k@1.0.0", max_cost=5) ``` ## Bring your own environment [Section titled “Bring your own environment”](#bring-your-own-environment) Your own tasks and reward are a Python module with two functions. `tasks(split)` lists the prompts of the `train` or `test` split with their answers, and `reward(completion, answer)` scores one completion: ```python def tasks(split): return [{"prompt": "Spell the word backwards...\n\nWord: valley", "answer": "yellav"}, ...] def reward(completion, answer): found = ANSWER.findall(completion) if not found: return None # no answer in the completion: InvalidOutput return 1.0 if found[-1].lower() == answer else 0.0 ``` A reward of 1 (or `True`) is `Correct`, anything lower `Incorrect`, and `None` `InvalidOutput`. If `reward` raises, the verdict is `InfrastructureFailure`, never a reward of 0. Answers reach only `reward`; the trainer and the model see the prompt and the optional `metadata`. A test prompt never appears in train, even if your lists repeat it. The quickest way in needs no Docker: publish the module as a pip package whose `load_environment()` returns an object (or the module itself) with `tasks` and `reward`, and name the exact version. Nodus builds it into an image when you apply the Environment; if the package does not load, the build fails and the Image `env--` shows the step and its log: ```yaml apiVersion: nodus.dev/v1 kind: Environment metadata: name: reverse-words spec: version: 1.0.0 package: pip: {name: reverse-words, version: 1.0.0} # index: for a private index loader: load_environment # the default; or module:function category: Custom modes: [Train, Evaluate] rewardType: Binary ``` ## Run an environment from the Environments Hub [Section titled “Run an environment from the Environments Hub”](#run-an-environment-from-the-environments-hub) An environment published on [Prime Intellect’s Environments Hub](https://app.primeintellect.ai/dashboard/environments) is a pip package, so it runs on Nodus as it is. Name it with the Hub’s index for its owner and the version the Hub lists: ```yaml apiVersion: nodus.dev/v1 kind: Environment metadata: name: reverse-text spec: version: 0.1.4 package: pip: name: reverse-text version: 0.1.4 index: https://hub.primeintellect.ai/primeintellect/simple/ category: Custom modes: [Train, Evaluate] ``` Its `load_environment()` returns a `verifiers` environment, and Nodus trains on it directly: the `dataset` is the train split, the `eval_dataset` the held-out one (with only one of them, every fifth task is held out), and each completion is scored by the environment’s own rubric, with `verifiers` 0.1 through 0.3. A multi-turn environment, such as a game like `wordle`, runs too: see [Train on a multi-turn environment](#train-on-a-multi-turn-environment), and so does one where the model calls tools: see [Train on a tool-calling environment](#train-on-a-tool-calling-environment). An environment whose tools run in a remote sandbox, a browser or an MCP server fails the build with that reason, as does one scored by another model (a `JudgeRubric`), since grading has no network and holds no credential for that model. The console marks each kind in its Hub search from the Hub’s own tags. The package comes from the Hub index alone and its dependencies may also come from PyPI. Grading has no network. Nodus fetches the datasets an environment loads while it builds the image, and every task list and grade after that reads that copy, so a run always sees the data its build saw. ## Train on a multi-turn environment [Section titled “Train on a multi-turn environment”](#train-on-a-multi-turn-environment) In a multi-turn environment the model and the environment take turns: the model replies, the environment answers with feedback, and the episode goes on until the environment ends it. Set `maxTurns` on the TrainingJob (**Turns per episode** under **Advanced** in the console) to let an episode run that many model turns: ```yaml spec: runtime: nodus/grpo-lora environment: {name: wordle@0.1.7, trainTasks: 512, heldOutTasks: 64, seed: 42} parameters: maxTurns: 6 # model turns per episode maxCompletionLength: 256 # tokens per reply maxEpisodeTokens: 2048 # tokens of the whole episode after the prompt, the model's and the environment's ``` The reward is the one the environment gives the whole episode, and training learns only from the model’s own tokens. An episode that reaches either cap ends there and is graded as it stands. With the default `maxTurns: 1`, only the first reply is graded. Baseline and final evaluation play the same episodes greedily. An [OpenEnv](https://github.com/meta-pytorch/OpenEnv) environment runs in process from a package whose `load_environment()` returns it. Each task is a seed: Nodus resets the environment with it, and every action is the model’s reply as the one text field of the environment’s action class. The environment must play the same episode for the same seed. Seeds 0 to 799 train and 800 to 999 are held out, unless the environment sets `train_seeds` and `test_seeds`: ```python import nltk from textarena_env.server.environment import TextArenaEnvironment def load_environment(): try: # the build downloads NLTK's word lists; grading reads them offline nltk.data.find("corpora/words") nltk.data.find("taggers/averaged_perceptron_tagger_eng") cached = True except LookupError: cached = False return TextArenaEnvironment("Wordle-v0", download_nltk=not cached) ``` A module of your own can also be multi-turn: give it `step(turns, answer)` in place of `reward`. It gets every model turn so far and returns the environment’s next message, or `None` once the episode ended, and the reward so far. Grading keeps no state between turns, so `step` replays the turns from the start. ## Train on a tool-calling environment [Section titled “Train on a tool-calling environment”](#train-on-a-tool-calling-environment) In a `verifiers` `ToolEnv` the model calls the environment’s tools and reads their results. Each task’s prompt offers the tools through the model’s chat template, the model calls one by writing a `{"name": "...", "arguments": {...}}` block (the form Qwen and most open chat templates teach), and the environment runs the call in its own code. The results are the next turn, and the episode ends when the model answers without calling a tool. Set `maxTurns` to the rounds of calls an episode may make: ```yaml spec: runtime: nodus/grpo-lora environment: {name: tool-test@0.1.1, trainTasks: 21, heldOutTasks: 43, seed: 42} parameters: maxTurns: 3 maxCompletionLength: 256 ``` The tools run inside the grader, which has no network, so a tool that searches the web or runs code in a remote sandbox cannot train. Pick a model whose chat template supports tools. ## Run any Reasoning Gym family [Section titled “Run any Reasoning Gym family”](#run-any-reasoning-gym-family) The catalog’s `nodus/reasoning-gym` serves five reviewed families. To train on any other [Reasoning Gym](https://github.com/open-thought/reasoning-gym) family, or a mix of them, publish a package whose `load_environment()` returns the dataset; its own `score_answer` scores each completion: ```python import reasoning_gym def load_environment(): return reasoning_gym.create_dataset("knights_knaves", size=2000, seed=7) ``` Every fifth entry is held out. The model’s last `…` is its answer, or the whole completion without one; partial credit is the reward and only a full score is `Correct`. Each family’s own licence applies. ## Run a verl or SkyRL dataset [Section titled “Run a verl or SkyRL dataset”](#run-a-verl-or-skyrl-dataset) A dataset prepared for [verl](https://github.com/volcengine/verl) or [SkyRL](https://github.com/NovaSky-AI/SkyRL) runs with its reward unchanged. Publish a module that names the parquet files with verl’s own keys, and your verl reward function under its usual name: ```python train_files = "hf://datasets/BytedTsinghua-SIA/DAPO-Math-17k/data/dapo-math-17k.parquet" val_files = "hf://datasets/BytedTsinghua-SIA/AIME-2024/data/aime-2024.parquet" # optional def compute_score(data_source, solution_str, ground_truth, extra_info=None): ... # verl's custom reward signature; a dict's "score" also works ``` Each row is verl’s: `prompt` (a system message and one user message at most), `data_source`, `reward_model.ground_truth` and `extra_info`. Without `compute_score`, each row’s `env_class` names the SkyRL-gym environment that scores it, built from the row as SkyRL builds it; add `skyrl-gym` to the package’s dependencies. Without `val_files`, every fifth row is held out. The package needs `datasets` as a dependency, and a SkyRL environment of more than one turn fails the build with that reason. To ship your own system packages or files, build the module into an image on the `env-base` image instead, push it and name its digest: ```dockerfile FROM ghcr.io/nodus-compute/catalog/env-base:1.0.0 COPY reverse_words.py /opt/environment/ ENV NODUS_ENVIRONMENT_LOADER=reverse_words NODUS_ENVIRONMENT=reverse-words NODUS_ENVIRONMENT_VERSION=1.0.0 RUN nodus-env info ENV HF_HUB_OFFLINE=1 HF_DATASETS_OFFLINE=1 ``` ```yaml apiVersion: nodus.dev/v1 kind: Environment metadata: name: reverse-words spec: version: 1.0.0 package: {image: registry.example.com/acme/reverse-words@sha256:...} category: Custom modes: [Train, Evaluate] rewardType: Binary ``` ```console $ nodus apply -f environment.yaml $ nodus get environment/reverse-words -w # Ready once the image digest is verified ``` A TrainingJob names it without the `nodus/` prefix (`environment: {name: reverse-words@1.0.0}`), and in Python `rl.grpo_lora(environment="reverse-words@1.0.0", ...)`. The whole example, with a GRPO TrainingJob, is in [`examples/training/custom-reward`](https://github.com/nodus-compute/nodus-platform/tree/main/examples/training/custom-reward). Nodus pulls the image from a public registry or from your organization’s space in the Nodus registry; other private registries are not supported for Environments yet. Any image works if it provides the two commands Nodus runs, as uid 10001 with no network: * `nodus-env tasks --split train|test --seed N` writes one JSON line per task: `{"taskId", "prompt", "metadata"}`. * `nodus-env grade` reads `{"taskId", "completion"}` lines and writes one `{"taskId", "verdict", "reward", "evidence"}` line for each, in order. The Environment becomes `Ready` once its image is verified, with the declared split sizes in `status.splits`. The first TrainingJob that uses a split and seed runs `nodus-env tasks` in one of its own grading Sandboxes and Nodus keeps that manifest for every later run of your organization, so the tasks never change between runs. A version’s image and graders cannot change: publish a new version instead, so earlier results stay reproducible. # Functions > Run Python functions on Nodus workers, deploy them as an App that stays up, and look them up from anywhere. A Function is a Python function that runs on Nodus workers. You decorate it, call it from your own code with `.remote()`, `.map()` or `.spawn()`, and Nodus starts workers when calls arrive, keeps them warm for a while and scales them back down. The pages of this guide cover [calling Functions](/docs/guides/functions/calls), [scaling them](/docs/guides/functions/scaling) and [what is billed while a worker waits](/docs/guides/functions/billing). ## Run an App [Section titled “Run an App”](#run-an-app) An App is the group of Functions in one file. `nodus run` creates an ephemeral App, runs your `main` on your machine and deletes the App when `main` returns. Every `.remote()` call runs on a Nodus worker. examples/functions/map/app.py ```python """One Function called three ways, then mapped over a thousand inputs. Run it with `nodus run examples/functions/map/app.py`. """ import nodus app = nodus.App("fn-map") @app.function(cpu=1, memory="1Gi", max_workers=4, target_concurrency=8, max_cost=1) def square(x: int) -> int: return x * x @app.function(cpu=1, memory="1Gi", max_cost=1) def divide(a: int, b: int) -> float: return a / b @app.local_entrypoint() def main(n: int = 1000) -> None: print("remote:", square.remote(7)) # one call; blocks for the result call = square.spawn(8) # starts a call and returns a handle print("spawned:", call.get(timeout=600)) results = list(square.map(range(n))) # one call per input, results in input order print("map:", len(results), "ordered:", results == [i * i for i in range(n)]) try: divide.remote(1, 0) except ZeroDivisionError as exc: # the remote exception comes back as its own type print("raised:", type(exc).__name__) ``` ```console $ nodus run examples/functions/map/app.py remote: 49 spawned: 64 map: 1000 ordered: True raised: ZeroDivisionError ``` The directory of the file (minus what `.gitignore` and `.nodusignore` exclude) is uploaded once per content hash, so the workers import the same code you ran. The App renews itself while `main` runs and is deleted shortly after `main` stops, even when your machine loses its connection. ## Deploy an App [Section titled “Deploy an App”](#deploy-an-app) `nodus deploy` keeps the App. Its Functions stay available after your program exits, and any program can call them. examples/functions/warm-pool/app.py ```python """A Function that keeps one worker warm, so a call never waits for a start. Deploy it with `nodus deploy examples/functions/warm-pool/app.py`. The idle worker is billed at the Function's worker rate for as long as it stays warm; set `min_workers=0` to pay only while calls run. """ import nodus app = nodus.App("fn-warm") @app.function(cpu=1, memory="1Gi", min_workers=1, max_workers=3, scaledown_window="2m", max_cost=1) def ping() -> str: return "pong" ``` ```console $ nodus deploy examples/functions/warm-pool/app.py ``` ```python import nodus ping = nodus.Function.from_name("fn-warm", "ping") print(ping.remote()) # pong ``` A deploy updates the App in place. A Function whose code, image or resources changed rolls its workers once, after their in-flight calls finish. A Function you removed from the file is deleted, and the calls it still had queued end as `Failed` with the reason `FunctionDeleted`, so no caller waits for a call nobody will run. ## What Nodus creates [Section titled “What Nodus creates”](#what-nodus-creates) | Object | What it is | Look at it with | | -------------- | ------------------------------------------------------------------------ | ------------------------------------------------------------ | | `App` | The Functions of one file; ephemeral for `run`, persistent for `deploy` | `nodus get apps` | | `Function` | One decorated function or class, with its image, resources and `scaling` | `nodus get functions` | | `FunctionCall` | One invocation, kept for seven days after it ends | `nodus get functioncalls -l nodus.dev/function=fn-warm-ping` | A Function reports its state in `status.phase`, how many workers it has in `status.workers`, how many calls wait in `status.queue`, and the expected start time in `status.estimate`. ```console $ nodus get function fn-warm-ping NAME PHASE WORKERS QUEUED COST AGE fn-warm-ping Running 1/3 0 $0.02 3m ``` ## Stop, restart and delete [Section titled “Stop, restart and delete”](#stop-restart-and-delete) ```console $ nodus stop function/fn-warm-ping # drain the workers; new calls wait in the queue $ nodus start function/fn-warm-ping # start workers again for the calls that waited $ nodus restart function/fn-warm-ping # replace every worker once, after in-flight calls finish $ nodus delete app/fn-warm ``` A stopped Function keeps accepting calls and holds them in the queue, so work you submit while it is stopped runs once it starts. Workers also stop by themselves when your balance cannot cover another renewal; calls queue until you add credit and then run. Note A call that runs for longer than its `timeout` (five minutes by default, 24 hours at most) ends as `Failed` with the reason `DeadlineExceeded`, and its worker restarts, because the thread that ran it cannot be interrupted. # What is billed while warm > How Function workers are billed when they run calls, when they wait, and when they scale to zero. You pay for worker time, by the second, at the worker’s rate. A call does not add a charge of its own. A Function shows the worker’s rate before you run anything, in `status.estimate.rateUSDPerHour` and in `f.estimate(...)`. ## When a worker is billed [Section titled “When a worker is billed”](#when-a-worker-is-billed) | State | Billed? | | -------------------------------------------------------------------- | ----------------------------------- | | A worker starting: placing, pulling its image, running `enter` hooks | Yes, from the moment it is placed. | | A worker running at least one call | Yes. | | An idle worker inside `min_workers` | Yes, always. This is warm capacity. | | An idle worker above `min_workers`, inside `scaledown_window` | Yes, until the window ends. | | A worker that was released | No. | | A Function with no workers | No. | So a Function with `min_workers=0` costs nothing while nothing is calling it, and each burst costs its workers’ time plus their `scaledown_window`. A Function with `min_workers=1` costs one worker’s rate every hour, and its calls never wait for a start. ## Reading the bill [Section titled “Reading the bill”](#reading-the-bill) Every worker is billed under one Function and shows up as two lines of usage: the time the worker spent running calls and the time it spent waiting (`warm_idle_seconds`). Both appear under the Function in your usage records, so the cost of keeping workers warm is never mixed into the calls themselves. Each call’s `status.costUSD` is the worker time attributed to that call: the worker’s rate divided by `target_concurrency`, for the whole seconds the call spent on a worker, at least one. A worker that runs four calls at once shows each of them a quarter of the rate. Call costs show where the busy time went. The bill is the workers’. ## Spending limits [Section titled “Spending limits”](#spending-limits) A Function does not take a spend cap yet: setting `max_cost` is refused when you deploy, so no cap can be set and silently ignored. Your credit balance is the limit on what its workers spend, and an idle `min_workers` worker bills until you stop the Function. When your balance cannot cover another renewal, the workers drain inside the reserve and the Function scales to zero. New calls queue and show `Funded=False` on the Function. Add credit and the Function starts workers again for the calls that waited. Nothing the calls were doing is lost: a call whose worker drained finishes first, and one that could not finish goes back to the queue. # Call a Function > Run one call, start calls without waiting, or map a Function over thousands of inputs, and read what comes back. A Function has three ways to start a call. Each call is a `FunctionCall` object, so you can look it up later, wait for it from another process or cancel it. | Call | What it does | | ------------- | -------------------------------------------------------------------------------------------------- | | `f.remote(x)` | Runs one call and blocks until it returns its result. | | `f.spawn(x)` | Starts one call and returns a handle: `handle.get(timeout=...)` waits, `handle.cancel()` stops it. | | `f.map(xs)` | Starts one call per input and yields the results in input order. | ```python with app.run(): print(square.remote(7)) # 49 handle = square.spawn(8) print(handle.get(timeout=600)) # 64 print(list(square.map(range(1000)))) # 1,000 results, in order ``` ## Map over many inputs [Section titled “Map over many inputs”](#map-over-many-inputs) `.map()` creates its calls in batches of up to 1,000 per request, and a batch is all or nothing: a request that fails leaves no calls behind, and sending it again does not create them twice. Results come back in input order whichever worker finishes first. `order_outputs=False` yields them as they finish, and `return_exceptions=True` yields a failed input’s exception instead of stopping the loop. ```python results = list(process.map(files, order_outputs=False, return_exceptions=True)) ``` Several callers can map into one Function at the same time. Workers take calls from each caller in turn, so a caller that submits 10 inputs is not stuck behind another caller’s 10,000. ## Arguments and results [Section titled “Arguments and results”](#arguments-and-results) Arguments and results are serialized with cloudpickle. Values up to 64 KiB travel inside the call. Larger ones are uploaded once, by content hash, and the call carries a reference, so a large argument costs nothing extra to send again. The caller and the worker image must run the same Python minor version, 3.10 to 3.13. ## Errors, retries and timeouts [Section titled “Errors, retries and timeouts”](#errors-retries-and-timeouts) An exception raised inside the Function is raised again in your process as its own type when that type can be imported there. Otherwise you get `nodus.errors.RemoteError`. Either way the remote traceback is attached. ```python try: divide.remote(1, 0) except ZeroDivisionError: ... ``` Two things can send a call back to the queue, and they are counted separately: * **The function raised.** With `retries=nodus.Retries(max_retries=3)` the call is retried with exponential backoff, up to ten times. Without retries the exception ends the call. * **The worker was lost.** A worker that disappears or is replaced does not count as a retry. Its calls go back to the queue and run on another worker, up to `recovery.maxAttempts` times (eight by default), then end as `Failed` with the reason `RecoveryLimitExceeded`. A call can finish only once. A result that arrives from a worker that no longer holds the call is refused, so a call that was re-dispatched never ends with two results. ## Look a call up later [Section titled “Look a call up later”](#look-a-call-up-later) ```python call = nodus.FunctionCall.from_name("fn-warm-ping-bcdfghjklm") print(call.get(timeout=60)) ``` ```console $ nodus get functioncalls -l nodus.dev/map= $ nodus get functioncall -o yaml # phase, result reference, retries, recoveries, costUSD ``` A call keeps its result for seven days after it ends. Each call reports `status.costUSD`, the worker time attributed to it; [the billing page](/docs/guides/functions/billing) explains how it relates to what you pay. # Scaling and warm workers > Set how many workers a Function keeps, how long an idle one stays, and how many calls each runs at once. A Function’s `scaling` bounds its worker pool. Nodus sizes the pool from the calls that are waiting and running. | Setting | Default | What it does | | -------------------- | ------- | ------------------------------------------------------------------------------------------ | | `min_workers` | 0 | Workers kept warm even when idle. [They are billed.](/docs/guides/functions/billing) | | `max_workers` | 10 | The most workers the Function ever has, 1 to 1,000. | | `scaledown_window` | `1m` | How long an idle worker above `min_workers` stays before it is released, up to 20 minutes. | | `target_concurrency` | 1 | Calls one worker runs at once, 1 to 1,000. Use more than 1 for I/O-bound functions. | ```python @app.function(cpu=2, memory="4Gi", min_workers=1, max_workers=20, scaledown_window="5m", target_concurrency=4) def embed(text: str) -> list[float]: ... ``` A pool with `target_concurrency=4` holds one worker for every four calls that are waiting or running, and never more than `max_workers`. A worker leaves only after it has been idle for the whole `scaledown_window`, so a burst that comes back inside the window finds its workers still there. ## Cold and warm starts [Section titled “Cold and warm starts”](#cold-and-warm-starts) A call that lands on a warm worker with a free slot starts at once. A call that needs a new worker waits for the worker to place, pull its image and start. Nodus shows both before you run anything: ```python print(embed.estimate("hello")) # expected cost, cold and warm start times, and the hold ``` ```console $ nodus get function embed -o jsonpath={.status.estimate} ``` `status.estimate.startup` holds the cold and warm bands, and `status.estimate.rateUSDPerHour` the worker’s rate. While workers are starting, the Function’s `Ready` condition says `WorkersStarting`. With no workers and nothing queued it says `ScaledToZero`, and with an image that is still building it says `ImagePending`. ## Classes: set up once per worker [Section titled “Classes: set up once per worker”](#classes-set-up-once-per-worker) A class keeps its state for the life of a worker. `@nodus.enter()` methods run once when the worker starts, before its first call. `@nodus.exit()` methods run once when the worker drains. Methods marked `@nodus.method()` get `.remote()`, `.map()` and `.spawn()`. examples/python/classes/app.py ```python """A class whose model loads once per worker, then serves many calls. Run it with `nodus run examples/python/classes/app.py`. """ import nodus app = nodus.App("classes") @app.cls(cpu=2, memory="4Gi", scaledown_window="5m", max_cost=1) class Greeter: @nodus.enter() def load(self) -> None: # Runs once when a worker starts, before its first call: load weights or open connections here. self.greeting = "hello" @nodus.method() def greet(self, name: str) -> str: return f"{self.greeting}, {name}" @nodus.exit() def close(self) -> None: self.greeting = "" @app.local_entrypoint() def main() -> None: greeter = Greeter() print(greeter.greet.remote("Ada")) print(list(greeter.greet.map(["Grace", "Linus"]))) ``` Load models and open connections in `enter`, so a call pays for them once per worker instead of once per call. A failing `exit` hook is logged and does not stop the others. ## Changing a Function [Section titled “Changing a Function”](#changing-a-function) Editing `scaling` takes effect on the next reconcile without restarting any worker. Changing the code, image, Python version or resources rolls the workers once: new workers start, and the old ones finish their calls and leave. `nodus restart` rolls them without a change. A Function whose image is still building keeps the workers it already has and keeps serving calls until the new image is ready. # Choose GPUs and check availability > See which accelerators are available right now, what they cost, and how to ask for exactly the hardware your run needs. You describe the hardware a run needs; Nodus finds capacity that fits and finishes it for the lowest expected cost. This guide shows how to see what is available, read an offering, and write a GPU request that says exactly what you mean. ## See what is available [Section titled “See what is available”](#see-what-is-available) ```console $ nodus get gpus NAME TYPE COUNT VCPU MEMORY REGION PRICE/H AVAILABILITY AGE a100-sxm-80g-x1-any A100-80G 1 30 200Gi unknown $1.27 Available 14s h100-sxm-80g-x8-us H100-SXM 8 208 1800Gi us $18.29 Available 14s l4-24g-x1-eu-int L4 1 8 32Gi eu $0.31 Limited 14s ``` `PRICE/H` is the from rate per machine-hour and `AGE` is how long ago the price and availability were confirmed. Filter with flags; each one narrows the list: ```console $ nodus get gpus --gpu H100 --count 8 --region us $ nodus get gpus --interruptible $ nodus get gpus --gpu L4 -o wide # typical and list rates, startup time and interruption rate ``` The console’s GPU picker shows the same list. ## Read an offering [Section titled “Read an offering”](#read-an-offering) An offering is a class of capacity: accelerator × count × region class, and whether it is interruptible. Its name spells that out: | Name | Meaning | | --------------------- | --------------------------------------------------------------- | | `h100-sxm-80g-x8-us` | Eight H100 SXM 80 GB GPUs per machine, in the `us` region class | | `l4-24g-x1-eu-int` | One L4 per machine in `eu`, interruptible | | `a100-sxm-80g-x1-any` | Capacity with no guaranteed location (`any`) | | `cpu-8c-32g-us` | An 8 vCPU, 32 GiB CPU machine in `us` | Each offering shows: * **From** and **typical** rates per hour. The from rate is that of the cheapest machine free now, the one a run starts on by default, or of the cheapest machine when none is free; the host shape shown is that machine’s. Your rate is shown in the estimate before launch and frozen for the run. No rate is ever above the list price. * **Availability**: `Available`, `Limited` (a few machines, or a count that is not reported) or `Unavailable`. * **Startup**: how long a machine usually takes to be ready (p50 and p90), from what Nodus has measured. * **Interruption rate** for interruptible offerings: how often such capacity is taken back per hour. Prices and availability refresh continuously. The public [pricing page](/pricing) shows the same from rates. ## Region classes [Section titled “Region classes”](#region-classes) `placement.regions` takes region classes, not data centers: `us`, `ca`, `eu`, `uk`, `in`, `apac`, `me` and `latam`. Capacity whose location is not reported has the class `any` in its name and never satisfies a non-empty `placement.regions`, so a run that must stay in a region never lands there. ```yaml spec: placement: regions: [eu, uk] ``` ## Ask for the hardware you need [Section titled “Ask for the hardware you need”](#ask-for-the-hardware-you-need) `resources.gpu.type` lists every accelerator you accept. Order does not matter: the scheduler picks among them by expected cost to finish. | You write | Nodus may use | | ------------------------------ | -------------------------------------------- | | `H100` | Any H100 variant: SXM, PCIe or NVL | | `[H100, H200]` | Any H100 or H200, never another family | | `H100-SXM` | Only the H100 SXM | | `H100!` or `exact: true` | Only the family’s primary variant (H100 SXM) | | `minMemory: 80Gi` with no type | Any accelerator with at least 80 GiB per GPU | Modal spellings work as written: `A100-40GB`, `A100-80GB`, `A10G`, `L40S`, `H100`, `H200`, `B200`, `T4`, `L4`. ```yaml spec: resources: gpu: {type: [H100, H200], count: 8} placement: interruptible: Allow # use interruptible capacity only when it is cheaper to finish maxRateUSDPerHour: "30.00" ``` `count` is per machine (1, 2, 4 or 8). For more GPUs than one machine holds, see [multi-node training](/docs/guides/multi-node/) (Beta). ## Use every GPU of one machine [Section titled “Use every GPU of one machine”](#use-every-gpu-of-one-machine) `--gpu A6000:2` (or `count: 2`) gives your container both GPUs of one machine, numbered 0 and 1. `torchrun` starts one process per GPU without flags, because Nodus sets `PET_NPROC_PER_NODE` to the GPU count; a `--nproc-per-node` you pass yourself wins. train.py ```python """A DDP smoke run on every GPU of one machine. torchrun starts one process per GPU (PET_NPROC_PER_NODE).""" import os import torch import torch.distributed as dist from torch.nn.parallel import DistributedDataParallel as DDP dist.init_process_group("nccl") local_rank = int(os.environ["LOCAL_RANK"]) device = torch.device("cuda", local_rank) torch.cuda.set_device(device) # Every process contributes 1, so the sum is the world size. one = torch.ones(1, device=device) dist.all_reduce(one) model = DDP(torch.nn.Linear(16, 1).to(device), device_ids=[local_rank]) opt = torch.optim.SGD(model.parameters(), lr=0.1) for step in range(20): x = torch.randn(32, 16, device=device) loss = (model(x) - x.sum(dim=1, keepdim=True)).pow(2).mean() opt.zero_grad() loss.backward() opt.step() if dist.get_rank() == 0: print(f"world_size={int(one.item())} gpus={torch.cuda.device_count()} " f"device={torch.cuda.get_device_name(device)} loss={loss.item():.4f}") dist.destroy_process_group() ``` run.sh ```sh nodus run --name torchrun-multi-gpu --gpu A6000:2 --max-cost 1.00 -- torchrun train.py ``` The run prints `world_size=2 gpus=2` from the first process. The processes reach each other over NCCL on the machine itself; `/dev/shm` is sized for it (half the container’s memory), so you need no `--shm-size` or `--ipc=host`. ## Check before you launch [Section titled “Check before you launch”](#check-before-you-launch) A dry run returns the estimate without starting anything: ```console $ nodus apply -f job.yaml --dry-run=server -o estimate ``` It shows the offering Nodus would use, the expected cost p50 and p90, the startup time, the first hold and the minimum charge, and how long the estimate stays valid. When nothing fits, it says why, for example `CapacityUnavailable`, `NoListPrice`, `ExceedsRemainingBudget` or `MissesDeadline`, with a fix. ## When a run waits in Queued [Section titled “When a run waits in Queued”](#when-a-run-waits-in-queued) A run with no matching capacity stays `Queued` and is placed as soon as capacity appears, until `placement.queueTimeout` (24 hours by default). `nodus describe` shows the offerings that were considered and why each was rejected. To start sooner, widen `resources.gpu.type` or `placement.regions`, allow interruptible capacity, or raise `placement.maxRateUSDPerHour`. How Nodus chooses among offerings is explained in [placement and scheduling profiles](/docs/concepts/supply-placement/). # Images > Run on the catalog images, build your own from a few steps or a Dockerfile, and use private registries. Every Job, Sandbox, Function and Agent runs in a container image. You can use a catalog image, any public or private registry image, or an Image that Nodus builds for you. When you submit work, Nodus resolves the image’s tag to a digest and records it, so a retry or a recovery runs exactly the same bytes. ## Catalog images [Section titled “Catalog images”](#catalog-images) | Image | Contents | Default for | | ------------------------------------------------- | ------------------------------------------------- | -------------------- | | `nodus/python:3.12` (also `3.10`, `3.11`, `3.13`) | Debian slim, Python, uv, git | Jobs without a GPU | | `nodus/pytorch:2.8-cuda12.8` | CUDA 12.8 runtime, Python 3.12, PyTorch 2.8 | Jobs with a GPU | | `nodus/agent-tools` | Python 3.12, Node 22, git, ripgrep, the Nodus SDK | Sandboxes and Agents | ```console nodus get images -n nodus nodus run --gpu L4 --image nodus/pytorch -- python -c "import torch; print(torch.cuda.get_device_name())" ``` Catalog images run as uid 1000 and include `/bin/sh`. ## Build availability [Section titled “Build availability”](#build-availability) Current hosted deployments support catalog images and prebuilt public or private registry images. Image builds from steps, Dockerfiles and Sandbox snapshots are unavailable. Those requests return `ImageBuildUnavailable` with HTTP 503 before accepting a build or charging for one. Publish your image to a registry, then use its reference directly or create an Image with only `base` and no steps. Environment packages that need a pip build must instead name a prebuilt `package.image`. Existing unstarted build requests report `Failed` with an explanation instead of waiting indefinitely. The build examples below describe the contract for a deployment with a configured build runner and registry. ## Build an Image [Section titled “Build an Image”](#build-an-image) An Image starts from a `base` and adds `steps`: `aptInstall`, `pipInstall`, `uvPipInstall`, `uvSync`, `run`, `copy`, `env` and `workdir`. image.yaml ```yaml apiVersion: nodus.dev/v1 kind: Image metadata: name: ex-images-build spec: base: nodus/python:3.12 steps: - aptInstall: [jq] - pipInstall: {packages: ["requests==2.32.5"]} - env: {APP_MODE: example} ``` ```console nodus apply -f image.yaml nodus wait image/ex-images-build --for=condition=Ready ``` The build log streams from `GET …/images/{name}/log?follow=true`; the Python SDK prints it while it builds. Run on it with `imageRef`: job.yaml ```yaml apiVersion: nodus.dev/v1 kind: Job metadata: name: ex-images-build spec: imageRef: {name: ex-images-build} # runs the digest the Image built, pinned at admission command: [python, -c, "import os, requests; print('requests', requests.__version__, os.environ['APP_MODE'])"] ``` In Python: ```python image = (nodus.Image.from_registry("nodus/pytorch:2.8-cuda12.8") .apt_install("git").pip_install("transformers==4.57.6", "peft") .env({"HF_HUB_ENABLE_HF_TRANSFER": "1"})) ``` Builds run on Nodus and are billed per vCPU-second at the CPU rate. An Image whose spec matches one your org has already built, on the same base digest, reuses that digest without building (`status.build.cached: true`). An Image with only a `base` is ready at once: it is that image’s digest. You can also build from a Dockerfile, inline or uploaded with its context: ```yaml spec: dockerfile: inline: | FROM nodus/python:3.12 RUN pip install polars==1.33.1 ``` Build-time credentials go in `buildSecrets`: they are mounted for the build only and never stored in a layer. ## Private registries [Section titled “Private registries”](#private-registries) Create a `Registry` Secret ([Secrets](/docs/guides/secrets/#private-registries)) and name it under `imagePullSecrets`, on the Job or Sandbox for `image`, or on the Image for a private `base`. When work runs on a provider that pulls containers itself, Nodus copies the image by digest into your org’s space in the Nodus registry first, so your registry credential never reaches the provider. Such providers need `/bin/sh` in the image; an image without it gets the `ProviderContainerNeedsShell` warning and runs elsewhere. ## Pull a built Image yourself [Section titled “Pull a built Image yourself”](#pull-a-built-image-yourself) A built Image lives in the Nodus registry under your org; `status.reference` holds its full reference. Log in with any username and an API key as the password, then pull it: ```bash REF=$(nodus get image ex-images-build -o jsonpath='{.status.reference}') echo "$NODUS_API_KEY" | docker login "${REF%%/*}" -u nodus --password-stdin docker pull "$REF" ``` Viewers, and keys without the `images:write` scope, can pull but not push. ## Status and errors [Section titled “Status and errors”](#status-and-errors) `nodus get image ex-images-build` shows `PHASE`, `DIGEST` and `SIZE`. A failed build has a reason: | Reason or error | Meaning | | ---------------------------------- | ---------------------------------------------------------------- | | `BaseNotFound` | The base does not exist, or its registry refused the pull Secret | | `StepFailed` | A step exited non-zero; the build log shows which | | `BuildTimeout` | The build ran past its time limit | | `ImageNotFound`, `ImagePullFailed` | A Job’s image could not be resolved when you submitted it | | `ImageNotReady` | A Job or Sandbox names an Image that has not finished building | # Inference > Call hosted models through an OpenAI-compatible API, route with Indra, and pay per token. Nodus inference is an OpenAI-compatible API over hosted open models. Point any OpenAI SDK at the Nodus base URL, use a Nodus API key, and pay per token from your prepaid credits. ## Sign in and create a key [Section titled “Sign in and create a key”](#sign-in-and-create-a-key) Sign in to the console, open **Settings → API keys**, and create a key with the `inference:invoke` scope. ## Call a model [Section titled “Call a model”](#call-a-model) ```python from openai import OpenAI client = OpenAI(base_url="https://inference.nodus-compute.ai/v1", api_key="YOUR_NODUS_API_KEY") reply = client.chat.completions.create( model="nodus/gpt-oss-120b", messages=[{"role": "user", "content": "What is the capital of France?"}], ) print(reply.choices[0].message.content) ``` ```bash curl https://inference.nodus-compute.ai/v1/chat/completions \ -H "Authorization: Bearer $NODUS_API_KEY" -H "Content-Type: application/json" \ -d '{"model": "nodus/indra", "stream": true, "messages": [{"role": "user", "content": "Hello"}]}' ``` Streaming works as in the OpenAI API (`stream: true`). Completed chat streams end with `data: [DONE]`; an interrupted or errored stream does not receive a generated completion marker. Every response carries a `Nodus-Request-Id` header. ## From the CLI [Section titled “From the CLI”](#from-the-cli) The CLI calls the same API with the credential of your current context: ```console $ nodus inference models MODEL NAME CONTEXT MAX OUTPUT PREVIEW nodus/indra Indra - - false openai/gpt-oss-120b GPT-OSS 120B 128K 64K false openai/gpt-oss-20b GPT-OSS 20B 128K 64K false $ nodus inference chat --model nodus/gpt-oss-20b --max-tokens 64 "What is the capital of France?" The capital of France is Paris. request ireq_01k6d2x7q9fvjt3y8m0c4r5n2e · openai/gpt-oss-20b · 84 input + 21 output tokens · $0.000015 $ nodus inference receipt ireq_01k6d2x7q9fvjt3y8m0c4r5n2e ``` `chat` prints the answer on standard output and the request id, the model that answered, the tokens and the charge on standard error; `-o json` prints the whole completion. `receipt` shows any request’s charge for 30 days. ## Models [Section titled “Models”](#models) `GET /v1/models` lists the models you can call now. Each model has a catalog name such as `openai/gpt-oss-120b` and a `nodus/` alias such as `nodus/gpt-oss-120b`. Models marked preview have provisional prices. In the console, open **Inference → Models** and select Indra to send a request to `nodus/indra`, or choose a model from the catalog. The response summary’s **Model** field shows the model identifier submitted with that request, including `nodus/indra` when using Indra. The API’s routing metadata and receipt still record which underlying model served the request. ## Responses, Messages and embeddings [Section titled “Responses, Messages and embeddings”](#responses-messages-and-embeddings) Chat models also accept stateless Responses and Anthropic Messages requests, with the same model ids, credits, endpoint limits and receipt headers. Responses supports text, inline images and client function tools; stored responses, background work and hosted tools are refused. Text Completions accepts one string prompt and one choice, without echo, suffix, logprobs or best-of sampling. Messages token counting is not available. ```python response = client.responses.create(model="nodus/gpt-oss-20b", input="Hello", max_output_tokens=64, store=False) print(response.output_text) vectors = client.embeddings.create(model="nodus/bge-m3", input=["first document", "second document"]) print(vectors.data[0].embedding) ``` Embeddings are available when the embedding model appears in `GET /v1/models`. They charge only input tokens. A model’s `operations` field identifies the routes it accepts, including `Embeddings`, and `aliases` lists its alternative model ids. The console labels embedding models as **text → vectors** and provides embedding examples in the model panel and endpoint API tab. A request can carry up to 2,048 inputs, each within the model’s context window. Messages accepts the same chat models at `/v1/messages`, with `max_tokens`, `messages` and optional `stream: true`; it does not expose private models used by managed agents. Streaming client tool arguments remain intact when several calls interleave. ## Audio [Section titled “Audio”](#audio) Transcribe or translate an audio file up to 25 MiB (flac, mp3, mp4, mpeg, m4a, ogg, wav or webm) with `openai/whisper-large-v3`, and generate speech with `canopylabs/orpheus-v1-english`: ```python with open("meeting.m4a", "rb") as f: text = client.audio.transcriptions.create(model="nodus/whisper-large-v3", file=f) print(text.text) speech = client.audio.speech.create(model="canopylabs/orpheus-v1-english", voice="tara", input="Your job finished.", response_format="wav") speech.write_to_file("done.wav") ``` Upload the file itself: audio URLs and streamed transcription are not supported. Transcription and translation are billed per audio hour for the file’s length, with a 10-second minimum per request; speech is billed per million input characters, and its audio is limited to 32 MiB. ## Indra: `nodus/indra` [Section titled “Indra: nodus/indra”](#indra-nodusindra) Send `model: "nodus/indra"` and Indra picks the catalog model best suited to each request: small, fast models for simple requests and stronger models for hard ones. The `X-Nodus-Routed-Model` response header names the model that answered. If routing is unavailable, Indra answers with its default model, so requests never fail because of routing. Indra was first called Composer, and `nodus/auto` still works as another name for it. Indra comes in two tiers, and the `X-Nodus-Composer-Tier` response header (`free` or `paid`) names the one that answered: * **Free Indra** is for every organization without the Indra plan. It routes among the three cheapest catalog models that are available (by blended price, three input tokens to one output token) and never draws on your credits. It is limited per organization to 100 requests and 200,000 tokens per UTC day, and each request to 4,096 output tokens (fewer when the day has less left); past either daily limit, requests get `429` with code `QuotaExceeded` until 00:00 UTC. Free Indra can also pause for the rest of the day when it is in heavy demand. * **Paid Indra** comes with the $20 per month Indra plan. It routes across every catalog model, and you pay the routed model’s per-token price plus the small routing charge, from the plan’s monthly Indra allowance first and then from your credits. Calling a model by name, directly or through a named endpoint, is always billed per token from your credits. ## Named endpoints and limits [Section titled “Named endpoints and limits”](#named-endpoints-and-limits) An `InferenceEndpoint` gives a model its own base URL and access policy: ```console $ nodus create inferenceendpoint support-bot --model nodus/gpt-oss-120b --rpm 120 --max-concurrent 8 --max-cost 50 $ nodus wait ep/support-bot --for=jsonpath={.status.phase}=Running $ nodus get ep NAME PHASE MODEL READY RPM CONCURRENT MAX-COST AGE support-bot Running nodus/gpt-oss-120b True 120 8 $50 12s $ nodus inference chat --endpoint support-bot "Where is my order?" ``` Call it at `https://inference.nodus-compute.ai/endpoints//support-bot/v1`, or send `model: "endpoint/support-bot"` on the shared base URL. Endpoints enforce `rpm`, `tpm`, `maxConcurrent`, `allowedKeys` (API key names, `--allowed-key`) and an optional `maxCostUSD`: once the requests through an endpoint have spent its cap, the next one gets `402 BudgetExceeded`, and raising `maxCostUSD` lets requests through again (it can only be raised). A new endpoint answers `503` for the few seconds until it is `Running`. `nodus stop ep/support-bot` makes it answer `503` until `nodus start ep/support-bot`; its `Ready` condition says whether its model can serve now. ## Billing per token [Section titled “Billing per token”](#billing-per-token) * Before a request runs, Nodus holds its maximum cost: the input bound plus `max_tokens` at the model’s rates. Lower `max_tokens` to hold less. * When it finishes, you are charged the tokens used (input, cached input and output) and the rest of the hold is released. A stream you stop early is charged for the tokens it produced. * Send an `Idempotency-Key` header to retry safely: a repeat within 24 hours returns the same answer with `Idempotent-Replayed: true` and no second charge. A request that failed without a charge (`Released`) frees its key, so the retry runs again. Reusing a key for a different request body returns `409`. * `GET /v1/requests/{id}` returns the receipt for 30 days: its `operation`, `usage`, `amountUSD`, `pricebookVersion` and `state`: `Running`, `Settled`, `Released` (no tokens, no charge) or `Unknown`. A request whose outcome Nodus never learned is `Unknown` and is never charged. ## Errors [Section titled “Errors”](#errors) Errors use the OpenAI shape, with the Nodus error code in both `type` and `code`: | Status | Code | Meaning | | -------- | --------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 400 | `Unsupported` | A feature outside per-token billing: built-in or server-side tools, a non-default `service_tier`, `store`, `background`, file parts, image URLs or audio URLs, a key repeated in another letter case, or an operation the model does not serve | | 400 | `Invalid` | A malformed request, or an upload that is not audio in a supported format | | 401 | `Unauthorized` | Missing or invalid API key | | 402 | `InsufficientCredits`, `BudgetExceeded` | Not enough credit for the hold, or a budget or endpoint cap blocks it | | 409 | `RequestInProgress` | The same `Idempotency-Key` is still running; retry after `Retry-After` | | 409 | `IdempotencyKeyReused` | The `Idempotency-Key` was used for a different request; send a new key | | 413 | `RequestEntityTooLarge` | A JSON body over 1 MiB or an audio file over 25 MiB | | 429 | `TooManyRequests` | An org or endpoint limit; retry after `Retry-After` | | 502, 503 | `Unavailable` | The model failed, is at capacity or is not available; retry shortly. You are not charged | # Jobs > Run a command to completion on a GPU or CPU, watch it, and download what it wrote. A Job runs your command until it finishes. You say what it needs (a GPU type, memory, a time limit, a spending cap) and Nodus runs it on capacity that fits, streams its logs, collects its outputs and charges only for the time it used. ## Sign in [Section titled “Sign in”](#sign-in) ```console $ pip install nodus-compute $ nodus login ``` `nodus login` opens the console in your browser and stores a key for this machine. ## Submit a Job [Section titled “Submit a Job”](#submit-a-job) The fastest way is `nodus run`. It uploads the current directory, starts the command and follows it until it exits, passing the exit code through: ```console $ nodus run --gpu L4 --image nodus/pytorch -- python hello.py ``` hello.py ```python import json import os import torch name = torch.cuda.get_device_name(0) print(f"Hello from {name}") # Anything written under /nodus/outputs is collected when the Job succeeds. report = {"gpu": name, "cuda": torch.version.cuda, "index": os.environ["NODUS_INDEX"]} with open(os.path.join(os.environ["NODUS_OUTPUTS_DIR"], "report.json"), "w") as f: json.dump(report, f) ``` Before anything runs, `nodus run` prints the estimate: the expected cost to completion and when it should start. Add `--max-cost 5` to stop the Job if it would spend more than $5, and `--timeout 2h` to limit its wall time. To keep a Job in version control, write it as a manifest and apply it: job.yaml ```yaml apiVersion: nodus.dev/v1 kind: Job metadata: name: hello spec: image: nodus/pytorch command: - python - -c - | import json, os, torch name = torch.cuda.get_device_name(0) print(f"Hello from {name}") with open(os.path.join(os.environ["NODUS_OUTPUTS_DIR"], "report.json"), "w") as f: json.dump({"gpu": name, "index": os.environ["NODUS_INDEX"]}, f) resources: gpu: L4 timeout: 10m maxCostUSD: "0.25" outputs: - name: report path: /nodus/outputs/report.json ``` ```console $ nodus apply -f job.yaml job.nodus.dev/hello created ``` Anything under `/nodus/outputs` is collected when the Job succeeds. Declare the paths you want to download by name under `outputs`. `kubectl apply -f job.yaml` works too once your kubeconfig points at Nodus. You do not have to name a GPU. Set `resources.gpu.minMemory` instead of `resources.gpu.type` and Nodus picks any accelerator with at least that much memory per GPU. If you also leave out `minMemory`, the annotations `nodus.dev/model` (for example `meta-llama/Llama-3.1-8B`) and `nodus.dev/dataset-bytes`, which you set under `metadata.annotations`, set it for you: 24Gi for 7B and 8B models, 32Gi for 13B and 14B, 48Gi for 34B and 40B, 80Gi for 70B and 72B, and 80Gi for a dataset over 1 TiB. `nodus apply -f job.yaml --dry-run=server -o yaml` shows the floor it picked. ## Watch it [Section titled “Watch it”](#watch-it) In the console, open **Runs**, then select a run. **Logs** shows reported progress, stored outputs and the latest recovery checkpoint alongside the log stream. Distributed and indexed runs can be narrowed to a worker rank or index. The stream shows one worker at a time, starting with rank 0 and index 0. It follows that worker’s latest attempt; it is not a history of every retry. Connection failures show an error and **Retry logs**, while **Download loaded logs** saves the lines currently loaded in the browser. **Run activity** shows scheduling and lifecycle events separately from program output. The command field preserves quoted arguments such as `python -c "print(1 + 1)"`. Shell pipelines and redirects need an explicit shell command, for example `bash -lc 'python prepare.py && python train.py'`. ```console $ nodus get job/hello -w $ nodus logs job/hello -f $ nodus describe job/hello ``` `get` shows the phase, the GPU, the attempt, the cost so far and the age (`-o wide` adds the offering, its rate and progress). `describe` adds the conditions, the events and the reason for a failure with a suggested fix. A Job moves through these phases: | Phase | What is happening | | ---------------------------------- | --------------------------------------------------------------------- | | `Queued` | Waiting for capacity that fits and for a funded hold | | `Provisioning` | Capacity is acquired; the image, source and inputs are being prepared | | `Running` | Your command is running | | `Recovering` | The capacity was lost; Nodus is moving the Job to new capacity | | `Suspending`, `Suspended` | Saving state and releasing compute, then paused | | `Cancelling` | Stopping and releasing compute | | `Succeeded`, `Failed`, `Cancelled` | Finished; nothing is running or billed | [Lifecycles](/docs/concepts/lifecycles/) lists every transition. ## Get the results [Section titled “Get the results”](#get-the-results) Open **Outputs** to download stored files and inspect their size and SHA-256. Distributed outputs identify their rank, and indexed outputs identify their index. A failed or cancelled run may have no stored outputs; its logs and activity explain what happened. **Details** includes attempts and the execution cost breakdown. ```console $ nodus cp job/hello:outputs/report ./report.json ./report.json: 214 bytes, sha256 4f1c0e9a2b7d ``` `nodus cp` checks every download against the SHA-256 recorded when the Job finished. Outputs stay available until you delete the Job. [Outputs](/docs/guides/outputs/) covers directories, Indexed Jobs and loading results into a database. ## Run many indexes [Section titled “Run many indexes”](#run-many-indexes) An Indexed Job runs the same command `completions` times, at most `parallelism` at once. Each run sees its number in `NODUS_INDEX` (and in `JOB_COMPLETION_INDEX`, as on Kubernetes): job.yaml ```yaml apiVersion: nodus.dev/v1 kind: Job metadata: name: indexed-outputs spec: image: python:3.12-slim # Four indexes, two at a time. Each index sees its number in NODUS_INDEX. completions: 4 parallelism: 2 command: - python - -c - | import json, os i = int(os.environ["NODUS_INDEX"]) rows = [{"index": i, "n": n, "square": n * n} for n in range(i * 100, (i + 1) * 100)] os.makedirs("/nodus/outputs/shards", exist_ok=True) with open(f"/nodus/outputs/shards/part-{i}.jsonl", "w") as f: f.writelines(json.dumps(r) + "\n" for r in rows) print(f"index {i}: wrote {len(rows)} rows") resources: cpu: "2" memory: 4Gi backoffLimit: 2 timeout: 15m maxCostUSD: "0.10" outputs: - name: shards path: /nodus/outputs/shards ``` The Job succeeds when every index has succeeded. `backoffLimit` is how many failed runs the whole Job tolerates before it fails; a retried index starts fresh on new capacity. `status.completedIndexes` lists the finished indexes, such as `0-2,5`. ## Limit time and cost [Section titled “Limit time and cost”](#limit-time-and-cost) | Field | Flag | What it does | | ------------------------ | --------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `maxCostUSD` | `--max-cost` | The most the Job may spend. At the cap it saves its state and becomes `Suspended` with reason `MaxCostReached`; raise the cap to resume it. The cap can only be raised. | | `timeout` | `--timeout` | Wall-clock limit counted from the first `Provisioning`; the Job fails with `DeadlineExceeded`. Time spent `Suspended` does not count. You can raise, lower or remove it on a running Job, and `status.timeoutTime` moves with it. `activeDeadlineSeconds` is accepted as an alias. | | `completeByTime` | `--complete-by` | When you need the result. Nodus picks capacity that should finish in time. | | `expectedDuration` | `--expected-duration` | Your estimate of the run time, used for the cost estimate before the Job has any history. | | `placement.queueTimeout` | | How long to wait for capacity before failing with `CapacityUnavailable`. The wait starts again when a suspended Job resumes. | When credits run out or a budget is reached, a Job saves its state and becomes `Suspended` with reason `InsufficientCredits` or `BudgetExceeded`. It stops by `status.stopByTime`, the time its funds run out, even if the save is not finished, and resumes on its own once it is funded again. ## Suspend, resume and cancel [Section titled “Suspend, resume and cancel”](#suspend-resume-and-cancel) ```console $ nodus suspend job/train # save state, release compute, stop billing $ nodus resume job/train # continue from the saved state $ nodus cancel job/train # stop for good ``` These set `spec.state` to `Suspended`, `Running` or `Cancelled`, so the same change works from a manifest. If a suspend cannot save state, the Job goes back to `Running` and its `Suspended` condition says `SuspendFailed`; nothing is lost. Nodus tries again ten minutes later, up to three tries in all (`status.suspendFailures` counts the failures), and then leaves the Job running. A Job with `recovery.continuity: Ephemeral` keeps no state, so suspending it restarts it from the beginning, and `nodus suspend` warns first. `nodus delete job/x` cancels a running Job before removing it. ## Recovery [Section titled “Recovery”](#recovery) Capacity can be lost while a Job runs, for example when interruptible capacity is reclaimed. What happens next depends on `recovery.continuity`: | Continuity | After a loss | | ------------------------ | -------------------------------------------------------------------------------------------- | | `Checkpointed` (default) | Resumes on new capacity from the last saved state in `NODUS_CHECKPOINT_DIR` (`/nodus/state`) | | `Restartable` | Starts again from the beginning on new capacity | | `Ephemeral` | Fails; nothing is kept | Your program writes and loads its own state files in `NODUS_CHECKPOINT_DIR`; Nodus saves and restores that directory. Set `recovery.onInterruption: Fail` to fail instead of recovering. An index that uses up `recovery.maxAttempts` fails the Job with `RecoveryLimitExceeded`. ## When a Job fails [Section titled “When a Job fails”](#when-a-job-fails) `status.reason` says why and `status.fix` says what to change. The common reasons: | Reason | Meaning | | ---------------------- | ------------------------------------------------------------------------------------------------ | | `BackoffLimitExceeded` | Your command exited non-zero more times than `backoffLimit` allows | | `OOMKilled` | The command ran out of memory; request more `resources.memory` or a larger GPU | | `DeadlineExceeded` | `timeout` elapsed | | `CapacityUnavailable` | No capacity fit the request within `placement.queueTimeout` | | `ImagePullFailed` | The image does not exist or its registry refused the pull | | `InvalidOutputPath` | A declared output was not written | | `NoProgress` | Two recoveries in a row each ran less than `recovery.minProgressDuration` before losing capacity | | `StorageUnavailable` | Nodus could not issue the storage credentials your command needs; run the Job again later | ## Clean up [Section titled “Clean up”](#clean-up) Finished Jobs are deleted `ttlSecondsAfterFinished` seconds after they finish. Jobs from `nodus run` default to 30 days; pass `--keep` to keep one. Deleting a Job deletes its outputs. ## Multi-node Jobs [Section titled “Multi-node Jobs”](#multi-node-jobs) Beta `spec.distributed` runs one Job across several nodes as a gang, with the rank, world size and rendezvous address in the environment. The Multi-node training guide covers launchers, networking and outputs from each rank. ## Environment [Section titled “Environment”](#environment) Every run sees these variables. Your own `env` cannot use the `NODUS_` prefix, except `NODUS_PARAM_`. | Variable | Value | | ------------------------------------- | ---------------------------------------------------------- | | `NODUS_JOB` | The Job’s name | | `NODUS_INDEX`, `JOB_COMPLETION_INDEX` | This run’s index, from 0 | | `NODUS_COMPLETIONS` | `spec.completions` | | `NODUS_OUTPUTS_DIR` | `/nodus/outputs` | | `NODUS_CHECKPOINT_DIR` | The first checkpoint path, `/nodus/state` by default | | `NODUS_INPUT_` | Where the input `` is mounted, under `/nodus/inputs` | | `NODUS_PARAM_` | A Sweep parameter (see [Sweeps](/docs/guides/sweeps/)) | ## Run code from GitHub [Section titled “Run code from GitHub”](#run-code-from-github) Install a GitHub Connection in the Job’s project and wait until it is Ready. Set `source.git.repo` to `owner/repository` or its GitHub HTTPS URL, and set `ref` to a branch, tag or commit. Nodus records the resolved commit and source archive in `status.source` before placing the Job. Retries and parallel indexes use the same files even when the branch moves. ```yaml source: git: repo: your-org/your-repository ref: main ``` The files appear in `workingDir` before the command starts. Set `source.git.path` to use another checkout directory; this does not change the command’s working directory. The checkout contains repository files, without `.git` history, submodule contents or Git LFS downloads. GitHub credentials stay outside the guest. The source archive is limited to 100 MiB compressed, 512 MiB expanded and 10,000 entries. A missing Connection or rejected archive prevents execution. # Logs and metrics > Follow a Job's output live, read it back after the capacity is gone, and watch GPU, CPU and memory use with nodus top. Everything your program writes to stdout and stderr is kept as its log. You can follow it while the program runs, read it back after it finishes, and filter it by time, attempt, index or rank. Usage samples (GPU, CPU and memory) are kept beside it, for `nodus top` and the console charts. ## Follow a run [Section titled “Follow a run”](#follow-a-run) `nodus run` streams the log until the Job finishes. To attach to a Job that is already running, or to come back after you detached: ```console $ nodus logs -f job/finetune-llama ``` `-f` (`--follow`) prints what is already stored, then new lines as they are written, without repeating or skipping any. Press Ctrl+C to stop following; the Job keeps running. The same command works for Sandboxes, Workspaces, Functions, Agents, AgentRuns, Images and individual Attempts: `nodus logs sb/dev`, `nodus logs attempt/NAME`. ## Choose which lines [Section titled “Choose which lines”](#choose-which-lines) | Flag | Shows | | ---------------------------- | ----------------------------------------------------------------------------------- | | `--tail 100` | The last 100 lines, then follows with `-f` | | `--since 10m` | Lines from the last 10 minutes (any duration: `30s`, `2h`) | | `--timestamps` | Each line prefixed with the time it was written | | `--attempt 2` | One attempt of the Job; the default is the latest one | | `--index 3`, `--all-indexes` | One index of an Indexed Job, or all of them | | `--rank 1`, `--all-ranks` | One rank of a multi-node Job (Beta), or all of them prefixed with `[r]` | | `--process 12` | The output of one Process in a Sandbox or Workspace, such as a `nodus exec` session | `--tail` and `--since` combine: `--since 1h --tail 50` prints at most the last 50 lines of the last hour. Without `--attempt`, a Job that was recovered shows its newest attempt; earlier attempts stay readable by number, which helps when a recovery followed a crash you want to look at. From Python, `job.logs()` returns the same text, and the MCP `logs` tool reads up to 2,000 lines at a time. ## Logs after the run [Section titled “Logs after the run”](#logs-after-the-run) Logs are stored as they are written, not only on the machine that ran the program. When a Job finishes, is suspended, or loses its capacity, `nodus logs job/NAME` still prints the whole log, and `nodus logs --attempt 1` reads an attempt whose machine is long gone. Logs are kept for 30 days. When capacity disappears without warning, the last few seconds written before the loss may be missing; everything before them is kept. Secret values you pass to the Job (`secrets:`) are replaced with `[redacted]` before the log is stored or shown. Logs are your program’s own output, exactly as it wrote it. Nodus’ own messages about scheduling and recovery are **Events** (`nodus describe job/NAME`, `nodus events`), never mixed into your log. `nodus events -w` prints the events already recorded and then each new one as it happens; add `--for job/NAME` to follow one object, and `-o json` to get the full Event objects. ## Watch usage with nodus top [Section titled “Watch usage with nodus top”](#watch-usage-with-nodus-top) `nodus top` shows what running work is using right now: ```console $ nodus top jobs NAME GPU GPU MEM CPU MEMORY $/H SPEND finetune-llama 94% 71.2Gi 3800m 41.0Gi $2.49 $6.12 eval-sweep-3 61% 18.5Gi 1200m 12.3Gi $0.80 $0.35 $ nodus top sandbox dev ``` GPU is the average utilization across the Job’s GPUs; GPU memory, CPU and memory are totals. Samples are taken every 15 to 30 seconds, so a value can be up to half a minute old. `nodus top sandboxes`, `workspaces` and `functions` work the same way. ## Charts and queries [Section titled “Charts and queries”](#charts-and-queries) The console’s **Metrics** tab charts the same samples over time. For your own dashboards and scripts, the history of one object is at: ```text GET /metrics/v1/namespaces/PROJECT/jobs/NAME/history?metric=gpu&since=6h&step=1m ``` `metric` is `gpu`, `gpuMemory`, `cpu` or `memory`, the resource can be `jobs`, `sandboxes`, `workspaces` or `functions`, and the answer has the shape of a Prometheus `query_range` result. Instead of `since`, pass `start` and `end` as RFC 3339 times. For anything else, the project’s samples answer PromQL in the Prometheus HTTP API at `/metrics/v1/namespaces/PROJECT/api/v1/query` and `query_range`, so Grafana and other Prometheus clients can use it as a data source with your API key. The series are: | Series | Labels | Value | | ----------------------------------------- | --------------------------------------- | ------------------------------------ | | `nodus_container_gpu_utilization_percent` | `kind`, `name`, `uid`, `attempt`, `gpu` | Utilization of one GPU, 0 to 100 | | `nodus_container_gpu_memory_bytes` | `kind`, `name`, `uid`, `attempt`, `gpu` | Memory in use on one GPU | | `nodus_container_cpu_millis` | `kind`, `name`, `uid`, `attempt` | CPU in use, in thousandths of a core | | `nodus_container_memory_bytes` | `kind`, `name`, `uid`, `attempt` | Memory in use | Every query is limited to the project in the path: a query that names another org or project is refused. For example, `max_over_time(nodus_container_gpu_utilization_percent{kind="Job",name="finetune-llama"}[1h])` gives the peak utilization of each GPU over the last hour. Training metrics such as loss and accuracy are separate: report them with `nodus.log.metrics(step=..., loss=...)` and they appear in the Job’s status under `status.progress.metrics`. # MCP > Let Claude Code, Cursor, Codex and other agents use Nodus through one MCP server, with confirmation before anything is created. Nodus runs one [MCP](https://modelcontextprotocol.io) server. An agent that speaks MCP can list your Jobs, read their logs, estimate what a manifest would cost, and, after you confirm, create it. The tools work on every kind of resource with the same names, so a new kind needs no new tools. The agent acts as you, with a role and scopes that you choose when you connect it. It can never do more than your own role allows. ## Connect a client [Section titled “Connect a client”](#connect-a-client) There are two ways to connect. The **hosted** server runs at `https://api.nodus-compute.ai/mcp` and signs you in with your browser. The **local** server is `nodus mcp`, which the client starts on your machine and which acts with your `nodus login`. Note `nodus mcp install CLIENT` writes the setup for you. It adds only a `nodus` entry to the client’s own configuration file, keeps a copy of the original next to it, and refuses to replace a different `nodus` entry unless you pass `--force`. Add `--dry-run` to see what would change. ### Claude Code [Section titled “Claude Code”](#claude-code) ```console $ claude mcp add --transport http nodus https://api.nodus-compute.ai/mcp ``` Run `/mcp` in Claude Code and choose **nodus** to sign in. Or for the local server: ```console $ nodus login $ nodus mcp install claude ``` ### Cursor [Section titled “Cursor”](#cursor) Use the Cursor link on the [connect page](/connect/), or: ```console $ nodus mcp install cursor --hosted ``` Drop `--hosted` to run `nodus mcp` locally instead. Cursor reads `~/.cursor/mcp.json`. ### Codex [Section titled “Codex”](#codex) ```console $ nodus mcp install codex --hosted $ codex mcp login nodus ``` Codex reads `config.toml` in `$CODEX_HOME` (`~/.codex` by default). Without `--hosted`, the entry starts `nodus mcp`. ### Other clients [Section titled “Other clients”](#other-clients) Any client that supports streamable HTTP can use the hosted URL. For one that starts a command, use [`/mcp.json`](/mcp.json) (the local server) or [`/mcp-hosted.json`](/mcp-hosted.json) (the hosted one). ## What the agent can do [Section titled “What the agent can do”](#what-the-agent-can-do) When you sign in to the hosted server, a consent page asks which org and which scopes the client may use. The tools it lists follow those scopes, so a client granted only `jobs:read` sees only the tools that read Jobs. | Tool | What it does | | ------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------- | | `whoami`, `api_resources`, `explain` | Who you are, the kinds and their fields. Agents start here. | | `get`, `describe`, `events` | Read objects, a summary with its recent events, and Events. | | `logs`, `wait`, `outputs` | Read logs (at most 2,000 lines), wait up to 60 seconds for a state, list outputs or get a download link valid for 10 minutes. | | `estimate` | Dry-run a manifest: the cost estimate, anything blocking it, and the `etag` that `apply` needs. | | `apply` | Create an object from a manifest. See below. | | `delete`, `set_state`, `request`, `record` | Delete an object, suspend or resume it, ask for an action such as a restart, and create a record such as a webhook replay. | | `exec`, `files` | Run a command in a running Sandbox, Workspace or Job (up to 300 seconds), and read, write or list its files. | A local server also has three tools that touch your machine’s files: `upload_source` (upload a directory as a code blob; it follows `.gitignore` and `.nodusignore`), `download_output` (save an output, verified against its checksum) and `cp` (copy a file to or from a Sandbox, Workspace or Job). ## Nothing is created without a confirmation [Section titled “Nothing is created without a confirmation”](#nothing-is-created-without-a-confirmation) `apply` does not create anything the first time. It returns the dry-run: the object as it would be stored, its estimated cost, and an `etag`. The agent shows you that. Only when you accept does it call `apply` again with `confirmed: true` and the same `etag`, and the create fails if the estimate changed in between. Objects created this way carry the label `nodus.dev/launched-by` with the value `mcp`. MCP never completes a payment. Applying a `TopUp` returns a checkout link for you to open yourself. ## Logs and output are data, not instructions [Section titled “Logs and output are data, not instructions”](#logs-and-output-are-data-not-instructions) Logs, command output and file contents come back marked as untrusted. They can contain anything a program wrote, so an agent must treat them as data. Nodus also stops a tool call that touches more than your grant allows, even if text in a log asks for it. ## Change or remove access [Section titled “Change or remove access”](#change-or-remove-access) Console › Settings › Connected agents lists each client with its scopes. Narrowing or deleting a grant takes effect within 30 seconds, even for a client that already holds a token. See [Connected agents](/docs/guides/access/connected-agents/). ```console $ nodus get oauthgrants $ nodus delete oauthgrant claude-code-3fa2c1 ``` ## Troubleshooting [Section titled “Troubleshooting”](#troubleshooting) | Symptom | What to do | | ------------------------------------------------------ | --------------------------------------------------------------------------------------------------------- | | The client lists fewer tools than you expect | The grant lacks the scope. Reconnect and approve more scopes, or check your role with `nodus auth can-i`. | | `not logged in` from `nodus mcp` | Run `nodus login`, or set `NODUS_API_KEY`. | | `install` says there is a different `nodus` connection | Remove that entry, or pass `--force` to replace it. | | `install` says it cannot parse the file | Fix the file’s syntax first. Nothing was changed. | | A write was refused with a changed estimate | Ask the agent to run `estimate` again and show you the new numbers. | # Migrate from Nodus Compute 0.x > What moves to the new Nodus automatically, what you recreate, and how the CLI, SDK, agents and billing change. Nodus 1.0 carries over your account, sign-in and organization memberships. Each migrated user receives $30 in free starter credit once, subject to email verification. Runs, Sandboxes and API keys must be recreated. ## What moves automatically [Section titled “What moves automatically”](#what-moves-automatically) | From 0.x | In Nodus 1.0 | | ----------------------------------------- | ------------------------------------------------------------------ | | Email/password, Google or GitHub sign-in | Your existing sign-in method | | Your account and organization memberships | An org with a `default` project and your mapped member role | | An account without an organization | Your sign-in is preserved; create your first org during onboarding | The old billing system used sandbox payments. Its balances, debt, saved cards, payments and invoice history do not carry over. Add a payment method in the new console when required to run work. ## Your $30 starter credit [Section titled “Your $30 starter credit”](#your-30-starter-credit) The grant is once per user across all organizations and expires 30 days after it is granted. If you already received starter credit in the new platform, migration does not grant it again, even if it has expired. Verified users with an imported organization receive the grant when that organization is released. If your email is not verified, verify it and open an organization you belong to; the console claims your credit automatically. If you have no organization, create one during onboarding. Your entitlement remains available while you complete these steps; its 30-day expiry starts only when the grant is issued. ## Cutover timeline [Section titled “Cutover timeline”](#cutover-timeline) Follow the announced cutover dates for pausing work and downloading outputs. The old service is kept available for the announced read-only window; running work and stored outputs are not copied into the new platform. ## What you recreate [Section titled “What you recreate”](#what-you-recreate) * **API keys.** 0.x keys (`nk_…`) stop working. Sign in with `nodus login` or create a new key in the console. * **Runs, Sandboxes and 0.x workspaces.** They are not copied. Download any files you need before the cutover date in your notice; the 0.x API keeps serving reads for 30 days afterwards. Recreate workspaces as `Workspace` resources. * **BYOC pools and hosts.** Recreate pools and enroll hosts in the new platform. * **Cloud accounts.** Re-connect AWS and GCP accounts as `CloudAccount` resources and remove the old role or grant. ## Update your tools [Section titled “Update your tools”](#update-your-tools) 1. **Upgrade the SDK and CLI.** `pip install -U nodus-compute` installs 1.0 under the same package, import and command names. The 1.0 API is new: a 0.x `Workload` becomes a `Job`, and `nodus run` replaces the 0.x submit commands. On first run, 1.0 moves an old `~/.nodus` config aside to `config.0x.bak`. 2. **Sign in again.** Run `nodus login`. `NODUS_BASE_URL` still works as an alias of `NODUS_API_URL` for this major version, with a warning. 3. **Reconnect coding agents.** Hosted MCP clients prompt you to authenticate again; local clients follow the setup at [/connect](/connect/). 4. **Update CI.** The 0.x GitHub Action fails with a migration message. Use `nodus-compute/run-action@v1` with a new API key. ## Billing changes [Section titled “Billing changes”](#billing-changes) Start-up and shutdown time is now billed 0.x started the meter when your command started. Nodus 1.0 bills rented machines from the moment they are created until their deletion is confirmed, at provider cost ÷ 0.875, and itemizes each charge as `Boot`, `Running`, `Restore` or `Teardown`. The same run can cost more than it did on 0.x, and the estimate shows each part before you launch. Links to 0.x docs pages redirect to their 1.0 equivalents, and old console links open a page that helps you find the object in the new console. # Multi-node training (Beta) > Run one training job across several machines with torchrun, Ray or your own launcher, and understand what it costs to assemble them. Beta Multi-node training is in Beta and on for every organization, with no access request and no purchase needed. Beta gangs have at most 8 nodes. When members on different providers connect over their public addresses (the `public` path), traffic between them, including gradients and the rendezvous, is **not encrypted**; each member accepts it only from the other members’ addresses. Keep `network: Colocated` if your data must not cross the internet unencrypted. A distributed Job runs your command on several machines at once, called a **gang**. Nodus acquires every member, connects them, checks that they can reach each other, and only then starts your command on all of them together. You write an ordinary training script; the launcher finds its peers from environment variables Nodus sets. ## Run a two-node job [Section titled “Run a two-node job”](#run-a-two-node-job) train.py ```python """A two-node DDP smoke run. torchrun reads its rendezvous from the PET_* variables Nodus sets on every rank.""" import os import torch import torch.distributed as dist from torch.nn.parallel import DistributedDataParallel as DDP backend = "nccl" if torch.cuda.is_available() else "gloo" dist.init_process_group(backend) local_rank = int(os.environ["LOCAL_RANK"]) device = torch.device("cuda", local_rank) if backend == "nccl" else torch.device("cpu") if backend == "nccl": torch.cuda.set_device(device) # Every process contributes 1, so the sum is the world size. one = torch.ones(1, device=device) dist.all_reduce(one) model = DDP(torch.nn.Linear(16, 1).to(device), device_ids=[local_rank] if backend == "nccl" else None) opt = torch.optim.SGD(model.parameters(), lr=0.1) for step in range(20): x = torch.randn(32, 16, device=device) loss = (model(x) - x.sum(dim=1, keepdim=True)).pow(2).mean() opt.zero_grad() loss.backward() opt.step() if dist.get_rank() == 0: print(f"world_size={int(one.item())} node={os.environ['NODUS_NODE_RANK']} " f"transport={os.environ['NODUS_GANG_TRANSPORT']} loss={loss.item():.4f}") dist.destroy_process_group() ``` run.sh ```sh nodus run --name torchrun-2-node --gpu H100 --nodes 2 --launcher torchrun --max-cost 2.00 -- torchrun train.py ``` `torchrun` needs no flags: its node count, rendezvous address and local address come from the `PET_*` variables on every rank. The run prints `world_size=2` from rank 0 when both nodes joined. ## Topology [Section titled “Topology”](#topology) Say how big the gang is with exactly one of: | Field (`nodus run` flag) | Meaning | | --------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------- | | `distributed.nodes` (`--nodes`) | The number of machines, 2 to 8 | | `distributed.totalGPUs` (`--total-gpus`) | The total GPU count; Nodus picks the machine shape. A total one machine can hold becomes a single-node Job | | `distributed.gpusPerNode` (`--gpus-per-node`) | GPUs per machine: 1, 2, 4 or 8. With `nodes` it defaults to the GPU count of `resources.gpu`; with `totalGPUs` Nodus picks it unless you set it | Every member gets the same GPU count. When `resources.gpu.type` lists several accelerator types, members may get different ones, each billed at its own rate. The resolved shape is written once to `status.topology` and never changes on a restart, so the world size stays the same for the whole run. A dry run shows it, with the hold for every member (`gangHoldUSD`) and the assembly bound: ```console $ nodus run --dry-run --gpu H100 --total-gpus 16 --launcher torchrun -- torchrun train.py ``` `distributed.network` bounds where members may be: | `network` | Members are placed | | --------------------- | --------------------------------------------------------------------------------------------------------------------- | | `Colocated` (default) | With one provider in one region, on its private network | | `Regional` | With any providers inside one region class | | `Global` | Anywhere. The run shows the `WANBound` condition, because training between distant machines is limited by the network | `distributed.transport` says which network paths you accept. `Direct` (default) admits every path that is not relayed: a provider’s private network, the members’ public addresses between providers (unencrypted in the Beta), and direct encrypted paths. `Auto` also admits **relayed** paths, which carry traffic through Nodus relays: they work between providers that cannot reach each other directly, but they are slow and fit small models only. The estimate warns `RelayedLowBandwidth` when a relayed path is possible. run.sh (across providers) ```sh nodus run --name torchrun-2-node-xp --gpu H100 --nodes 2 --launcher torchrun \ --network global --transport auto --startup-timeout 15m --max-cost 2.00 -- torchrun train.py ``` ## Launchers [Section titled “Launchers”](#launchers) | `launcher` | What each rank runs | | ------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `Plain` (default) | Your command once per node. `RANK`, `LOCAL_RANK=0`, `LOCAL_WORLD_SIZE=1` and `WORLD_SIZE` (the node count) make `env://` initialization work for one process per node | | `Torchrun` | Your `torchrun` command on every node; rank 0 hosts the rendezvous. Nodus owns restarts, so `PET_MAX_RESTARTS=0` | | `Ray` | A Ray cluster: rank 0 is the head, the others join it, and your command runs on rank 0 once all nodes are up, with `RAY_ADDRESS` set | | `Verl` | The Ray launcher plus `NODUS_VERL_OVERRIDES` (`trainer.nnodes`, `trainer.n_gpus_per_node`) to append to your verl command | | `Accelerate`, `Deepspeed` | The environment contract plus `ACCELERATE_*` or `/etc/nodus/hostfile`. Accepted in the Beta, qualified later | Every rank gets the same variables except the rank-specific ones: | Variable | Value | | -------------------------------------------------- | -------------------------------------------------------------------------- | | `NODUS_NODE_RANK`, `NODE_RANK` | The Nodus rank of this node; rank 0 hosts the rendezvous | | `NODUS_NUM_NODES`, `NNODES`, `NODUS_GPUS_PER_NODE` | The node count and GPUs per node | | `NODUS_NODE_IPS` | Every member’s address in rank order | | `MASTER_ADDR`, `MASTER_PORT` | Rank 0’s address and `29500` | | `NODUS_GANG_EPOCH`, `NODUS_GANG_TRANSPORT` | The current epoch and its path: `private`, `public`, `direct` or `relayed` | | `NODUS_RESTORE_URI` | The checkpoint to resume from after a restart, when one exists | `NODUS_NODE_RANK` is Nodus’s rank. torch assigns its own global `RANK` inside torchrun and may order nodes differently. Setting any `NODUS_*`, `PET_*` or `MASTER_*` variable, `NODE_RANK` or `NNODES` in your spec is rejected. Your `NCCL_*` settings are kept, except the few the network path decides. ## Failures and restarts [Section titled “Failures and restarts”](#failures-and-restarts) If a member’s machine is lost or reclaimed, Nodus stops the whole gang’s current epoch at once and restarts it: surviving machines are kept, only the lost ranks get new machines, and every rank starts again at the next epoch from the latest gang checkpoint (`NODUS_RESTORE_URI`). A process left over from the old epoch cannot join the new rendezvous. `recovery.maxAttempts` counts these restarts (default 3). The first failure’s cause is the gang’s reason, so a crash on one rank that brings down the others reports that crash. If a machine is refused or never appears while the gang is still assembling, only that rank gets another machine (up to two per assembly) and the registered ranks keep waiting at the start barrier. A suspend or cancel in progress is never turned into a restart: a machine lost during it simply completes the stop. If your command exits non-zero on any rank, the run fails; a zero exit on every rank succeeds it. Each rank’s attempts carry the `nodus.dev/rank` label, so `nodus get attempts -l nodus.dev/job=,nodus.dev/rank=1` lists rank 1’s attempts across restarts, with each attempt’s placement, boot or restore time and cost. ## The assembly bound [Section titled “The assembly bound”](#the-assembly-bound) Assembly waits at most `distributed.startupTimeout` (default 15 minutes, 5 to 60) for every member to become ready and pass the network check. If it does not, every acquired member is released, and Nodus tries again up to `maxAssemblyRetries` times (default 2) before the run fails with `GangAssemblyTimeout`. You pay for each member from the moment its provider starts billing until it is confirmed deleted, including time spent waiting for the other members. The estimate shows the most a failed assembly can cost, the **assembly bound**: the sum of the members’ rates plus up to two replacement machines, times the startup timeout plus teardown, times the number of tries. Nodus pays, not you, when assembly fails for a Nodus reason: a network check that fails on a path Nodus lists as qualified, an address collision, or an outage of the Nodus mesh. ## Beta caveats [Section titled “Beta caveats”](#beta-caveats) * Public paths carry traffic between members’ public addresses in clear text. Only the gang’s own members may connect to a member’s rendezvous and collective ports, but the bytes are not encrypted on the way. * Relayed paths carry every byte through a Nodus relay, are billed per relayed GiB, and suit small models only. * Relayed gangs have at most 2 members until larger relayed gangs are qualified. * On relayed paths your image must use dynamically linked glibc programs; otherwise the run fails with `ShimNotLoaded`. * Provider pairs are added as each is qualified; a pair that is not qualified is never offered. The measured TCP throughput and round-trip time per path class are published here as pairs are qualified. # Notifications > Which emails Nodus sends, who receives them, how often, and how to switch off the optional ones. Nodus emails you when something needs your attention: a balance running low, a Budget threshold, a Job that stopped because of money, a pool host that went offline. Each notice is sent once for the condition that caused it. A retry or a repeat of the same condition never sends a second email. Every notice is also an Event, and most are [webhook](/docs/guides/webhooks/) types, so a program can react to them too. ## Who receives what [Section titled “Who receives what”](#who-receives-what) | Kind of notice | Goes to | | -------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------- | | Money: balance, funding, Budgets, grants, top-ups, arrears, disputes, storage and egress quotas | The billing email of the org’s BillingAccount, and every Owner and Admin, once each | | Budget notices | The same people, and the addresses in the Budget’s `spec.notify.emails` (up to 10) | | About an object: for example a Job stopped by its cost cap | The person who created it. For an object created by an API key or service account, the key’s owner or the project’s Admins | | Pools: host offline, capacity exhausted, burst approval needed, forecast shortfall, an action proposed or executed, a cloud sync that failed | The org’s Owners and Admins, and the person who created the Pool | | Invites | The invited address, even before it belongs to a member | ## What Nodus sends [Section titled “What Nodus sends”](#what-nodus-sends) | Notice | When | How often | | ----------------------------------- | ------------------------------------------------------------------------ | -------------------------------------------- | | Low balance | Available credit fell below your warning level | At most once a day | | Funding lost | An object stopped because credit ran out | One digest per hour | | Maximum cost reached | An object stopped at its `maxCostUSD` | Once per object | | Budget threshold, Budget exceeded | A Budget reached a percentage, or refused a new run | Once per threshold and period | | Grant expiring | A credit grant expires in 7 days and in 1 day | Once each | | Top-up failed, auto-recharge failed | A charge did not go through | Once per attempt | | Arrears | A balance went negative | At most once a day | | Dispute opened or closed | A payment dispute changed | Once per change | | Storage or egress quota reached | A quota was hit | At most once a day | | Storage in arrears | Storage is unpaid after 7, 21 and 28 days, with the purge proposed at 30 | Once each | | Pool notices | A pool event listed above | Once per event and window | | Invite | Someone invited you to an org | Once per invite, and again when it is resent | Stripe sends the receipt for a top-up to your billing email. Sign-in and password emails come from your account system, not from these notices. ## Switch off the optional ones [Section titled “Switch off the optional ones”](#switch-off-the-optional-ones) Budget thresholds, low-balance warnings and run-completion emails are optional: each member can turn them off for themselves. Stops, disputes, arrears, quota limits and security notices are not optional. Change your preferences in your account settings. They are also at `GET` and `PUT /auth/me/settings`; read the current settings first to see the fields. Your choice applies to you only. Other recipients of the same notice still get it. ## Add recipients to a Budget [Section titled “Add recipients to a Budget”](#add-recipients-to-a-budget) ```yaml apiVersion: nodus.dev/v1 kind: Budget metadata: name: research spec: limitUSD: "500" period: Monthly notify: emails: [finance@example.com] ``` Note Notices never contain secrets. An invite link is sealed until the email is sent and is erased afterward. # Outputs > Collect files from a Job, download them with a verified checksum, and load results into Postgres. An output is a file or directory a Job writes for you to download. Outputs are collected when the Job succeeds, checksummed, and kept until you delete the Job. ## Write outputs [Section titled “Write outputs”](#write-outputs) Write anything you want to keep under `/nodus/outputs` (also in `NODUS_OUTPUTS_DIR`). Every file there is collected when the Job succeeds. To give a file or directory a name of your own, declare it: job.yaml ```yaml apiVersion: nodus.dev/v1 kind: Job metadata: name: hello spec: image: nodus/pytorch command: - python - -c - | import json, os, torch name = torch.cuda.get_device_name(0) print(f"Hello from {name}") with open(os.path.join(os.environ["NODUS_OUTPUTS_DIR"], "report.json"), "w") as f: json.dump({"gpu": name, "index": os.environ["NODUS_INDEX"]}, f) resources: gpu: L4 timeout: 10m maxCostUSD: "0.25" outputs: - name: report path: /nodus/outputs/report.json ``` A declared `path` is a file or a directory under `/nodus/outputs`. Every file is collected as its own output, named by its path below `/nodus/outputs`, and a declared name is another name for the file or directory it points at. A Job can declare up to 32 outputs; names use lowercase letters, digits, `.`, `_` and `-`, and the name `outputs` and the `nodus.` prefix are reserved. Outputs are separate from checkpoints: files in `NODUS_CHECKPOINT_DIR` are for resuming the Job, not for download. ## Download outputs [Section titled “Download outputs”](#download-outputs) ```console $ nodus cp job/hello:outputs/report ./report.json ./report.json: 214 bytes, sha256 4f1c0e9a2b7d $ nodus get job/hello -o jsonpath='{.status.outputs}' ``` `nodus cp` writes the file only after its SHA-256 matches the digest recorded when the output was collected. Name a file by its declared name (`report`) or by its path below `/nodus/outputs` (`report.json`); the files of a directory output are listed one by one in `status.outputs`. For an Indexed Job, every index writes its own copy of each output. Pick one with `--index`: ```console $ nodus cp job/indexed-outputs:outputs/shards/part-2.jsonl ./part-2.jsonl --index 2 ``` ### With the API [Section titled “With the API”](#with-the-api) `GET /apis/nodus.dev/v1/namespaces//jobs//outputs` lists the collected outputs with their index, size and `sha256`. `GET …/outputs/` answers `302` with a short-lived download URL and the digest in the `X-Nodus-SHA256` header; add `?index=` when several indexes produced the output. Verify the digest of what you download. Errors are `Status` objects with a `reason` and a `fix`: `OutputNotFound` (404) when the Job committed no such output, `OutputIndexRequired` (400) when you need to pass `?index=`, and `Invalid` (422) for a malformed `index`. ## Stage outputs [Section titled “Stage outputs”](#stage-outputs) In a [Pipeline](/docs/guides/pipelines/), a stage reads an earlier stage’s output as an input, mounted read-only at `/nodus/inputs/` and named by `NODUS_INPUT_`. Any Job can do the same with an `output` input that names another Job: ```yaml inputs: - name: data output: {job: prepare, name: shards} ``` ## Load outputs into Postgres [Section titled “Load outputs into Postgres”](#load-outputs-into-postgres) An output with a `sink` is loaded into a table in your database after the Job succeeds. The database is reached through a Postgres, Neon or Supabase Connection: job.yaml ```yaml apiVersion: nodus.dev/v1 kind: Job metadata: name: output-sink spec: image: python:3.12-slim command: - python - -c - | import json with open("/nodus/outputs/squares.jsonl", "w") as f: for n in range(1000): f.write(json.dumps({"n": n, "square": n * n}) + "\n") timeout: 10m maxCostUSD: "0.05" outputs: - name: squares path: /nodus/outputs/squares.jsonl # After the Job succeeds, the rows are loaded into this table through the Connection named analytics. sink: connection: analytics table: squares mode: Replace ``` * The output must be a `.csv`, `.jsonl` or `.parquet` file. * `mode: Append` (the default) adds rows; `Replace` replaces the table’s rows. * Each file can be up to 5 GB and 50 million rows, and each record up to 8 MiB. * The load runs as part of the Job and is billed to it as CPU time. `status.outputs[].sink` shows each load’s phase (`Pending`, `Loading`, `Loaded` or `Failed`) and the rows loaded, and the `SinksLoaded` condition turns true when every load has finished. To retry the failed loads of a finished Job: ```console $ nodus request reload-sinks job/output-sink ``` ## How long outputs last [Section titled “How long outputs last”](#how-long-outputs-last) Outputs stay downloadable until the Job is deleted, and deleting the Job deletes them. Jobs started with `nodus run` are deleted 30 days after they finish unless you pass `--keep`; set `ttlSecondsAfterFinished` on a manifest to choose your own retention. # Pipelines > Chain Jobs into stages that pass outputs along, run in parallel where they can, and share one spending cap. A Pipeline runs Jobs as stages. A stage starts when the stages it depends on have succeeded, reads their outputs as inputs, and runs in parallel with every other stage that is ready. The whole Pipeline shares one spending cap. ## Submit a Pipeline [Section titled “Submit a Pipeline”](#submit-a-pipeline) This Pipeline prepares a dataset on a CPU, trains on a GPU, and evaluates the trained model: pipeline.yaml ```yaml apiVersion: nodus.dev/v1 kind: Pipeline metadata: name: train-eval spec: # One cap for the whole run: every stage's spend counts against it. maxCostUSD: "0.50" failurePolicy: FailFast # Stage templates inherit these fields unless they set their own. defaults: image: nodus/pytorch timeout: 15m stages: - name: prepare template: resources: cpu: "2" memory: 4Gi command: - python - -c - | import os, torch os.makedirs("/nodus/outputs/shards", exist_ok=True) x = torch.randn(4096, 16) y = (x.sum(dim=1, keepdim=True) > 0).float() torch.save({"x": x, "y": y}, "/nodus/outputs/shards/data.pt") outputs: - name: shards path: /nodus/outputs/shards - name: train dependsOn: [prepare] inputs: # Mounted at /nodus/inputs/data and named by NODUS_INPUT_DATA. - name: data fromStage: prepare output: shards template: resources: gpu: L4 command: - python - -c - | import os, torch d = torch.load(os.path.join(os.environ["NODUS_INPUT_DATA"], "data.pt")) model = torch.nn.Linear(16, 1).cuda() opt = torch.optim.SGD(model.parameters(), lr=0.1) x, y = d["x"].cuda(), d["y"].cuda() for step in range(200): opt.zero_grad() loss = torch.nn.functional.binary_cross_entropy_with_logits(model(x), y) loss.backward() opt.step() print(f"final loss {loss.item():.4f}") os.makedirs("/nodus/outputs/model", exist_ok=True) torch.save(model.state_dict(), "/nodus/outputs/model/model.pt") outputs: - name: model path: /nodus/outputs/model - name: eval dependsOn: [train] inputs: - name: model fromStage: train output: model template: resources: cpu: "2" memory: 4Gi command: - python - -c - | import os, torch model = torch.nn.Linear(16, 1) model.load_state_dict(torch.load(os.path.join(os.environ["NODUS_INPUT_MODEL"], "model.pt"))) x = torch.randn(1024, 16) acc = ((model(x) > 0).float() == (x.sum(dim=1, keepdim=True) > 0).float()).float().mean() print(f"accuracy {acc.item():.3f}") ``` ```console $ nodus apply -f pipeline.yaml pipeline.nodus.dev/train-eval created ``` The console and the Python SDK also offer this prepare, train and evaluate shape as the train-eval template. ## Watch it [Section titled “Watch it”](#watch-it) ```console $ nodus get pipeline/train-eval -w NAME PHASE STAGES COST AGE train-eval Running 1/3 $0.02 3m $ nodus get jobs -l nodus.dev/pipeline=train-eval $ nodus logs job/train-eval-train -f ``` Each stage runs as a Job named `-`, so every Job command works on a stage: `logs`, `exec`, `describe` and `cp`. `status.stages` shows each stage’s phase, Job, start and finish times and cost. The pipeline name plus the stage name can be at most 62 characters. ## Stages and dependencies [Section titled “Stages and dependencies”](#stages-and-dependencies) * `dependsOn` lists the stages that must succeed first. The stages form a graph with no cycles; a cycle is rejected with the path that closes it. * `inputs` hands an output of an earlier stage to this one: `fromStage` names the stage (it must also be in `dependsOn`) and `output` names one of its declared outputs. The output is mounted read-only at `/nodus/inputs/`, and `NODUS_INPUT_` holds that path. * `maxParallel` limits how many stages run at once. It defaults to the number of stages, so every ready stage starts. * `defaults` holds Job fields every stage inherits. A field a stage sets wins; objects merge field by field, and a list a stage sets replaces the default list. ## When a stage fails [Section titled “When a stage fails”](#when-a-stage-fails) | `failurePolicy` | What happens | | -------------------- | ------------------------------------------------------------------------------------------------------------------------------- | | `FailFast` (default) | The running stages are cancelled, the stages not yet started are skipped, and the Pipeline fails | | `RunIndependent` | Stages that do not depend on the failed stage run to the end; the stages that depend on it are skipped; the Pipeline then fails | The Pipeline’s `status.reason` is `StageFailed` and its message names the failed stages. Each stage keeps its own `status.reason`, so `nodus describe job/-` shows why it failed. ## Cost cap [Section titled “Cost cap”](#cost-cap) `maxCostUSD` caps the whole Pipeline: the spend of every stage counts against it. When the Pipeline reaches the cap, every running stage saves its state and suspends, no new stage starts, and the Pipeline becomes `Suspended` with reason `MaxCostReached`. Raise the cap to resume: ```console $ nodus patch pipeline/train-eval --type merge --patch '{"spec":{"maxCostUSD":"1.00"}}' ``` The cap can only be raised. A stage can set a lower `maxCostUSD` of its own. ## Suspend, resume and cancel [Section titled “Suspend, resume and cancel”](#suspend-resume-and-cancel) `nodus suspend pipeline/x`, `resume` and `cancel` apply to every stage that has not finished. They set the Pipeline’s `spec.state`, which its stage Jobs follow. Deleting a Pipeline deletes its stage Jobs and their outputs; `ttlSecondsAfterFinished` deletes a finished Pipeline for you. # Run jobs on your own machines > Enroll your hosts into a pool, route jobs to them first, forecast demand and automate routine actions. A pool is a group of your own machines that Nodus schedules work onto before it rents anything. Nodus does not charge a compute rental fee for your hosts; you continue paying your cloud or hardware costs. Routing a job to a pool costs $0.02 per GPU-hour while a Nodus-scheduled attempt uses a GPU. CPU-only pool work has no device-hour fee. Predict capacity forecasts cost $99.00 per pool per month. You agree to each price when you turn it on. Pools have four parts: * **Measure**: enroll hosts with a one-line installer and see their inventory, utilization and health. * **Route**: send jobs to the pool first, wait for it or burst to the market under rules you set. * **Predict**: forecasts of demand and free capacity, and recommendations for sizing and wait policies. * **Act**: routine actions such as reclaiming idle hosts or draining in a maintenance window, with approvals. ## Create a pool [Section titled “Create a pool”](#create-a-pool) ```sh nodus create pool lab --routing prefer ``` The CLI and the console ask you to agree to the routing fee before the pool is created. In the console, open **BYOCompute** and choose **New pool**. ## Enroll a host [Section titled “Enroll a host”](#enroll-a-host) A host needs Linux on x86-64 or arm64 with a running systemd service manager and cgroup v2 with CPU and memory controllers. It also needs running containerd at `/run/containerd/containerd.sock`, a healthy `overlayfs` snapshotter, and `ctr`, `runc`, `containerd-shim-runc-v2` and `timeout` on its system path. Follow the [containerd setup instructions](https://github.com/containerd/containerd/blob/main/docs/getting-started.md) if these are missing. GPU hosts also need working NVIDIA drivers and the [NVIDIA Container Toolkit with CDI](https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/cdi-support.html). The installer checks these prerequisites before downloading the agent or consuming the enrollment token. It does not install system packages, reconfigure Docker or restart containerd. Download the installer from the URL in the enrollment command to `install.sh`, then check the host without enrolling it: ```sh sudo sh ./install.sh --check ``` A successful check confirms runtime prerequisites, not a completed GPU Job. Create an enrollment token for the pool; the response shows the installer once: ```sh nodus create enrollmenttoken --pool lab ``` `--ttl 2h` shortens the token’s life and `--label rack=a` copies a label onto the node when it enrolls. The token needs no name; the server gives it one. Run the command it prints as a user who can `sudo` on the host. It writes the token to a file only you can read, downloads the signed `nodusd` agent, checks its sha256 checksum, installs it as a systemd service and joins the pool. It also verifies the signature when `cosign` is installed; `--require-signature` makes that check mandatory. The token works once and expires after 24 hours; the console shows the countdown. Wait for the node to report `Ready` before submitting work. The installer is for enrollment; upgrading an existing host requires draining its work, replacing the verified agent binary and restarting `nodusd`. ```sh nodus get nodes nodus get node/host-1 -o yaml ``` Each job runs in a network namespace of its own, and `nodusd` keeps its network rules in the nftables table `inet nodus`. Outside that table it adds two rules to the host’s firewall, `-i ndv+ -j ACCEPT` and `-o ndv+ -j ACCEPT`. They match only its job interfaces, because Docker and ufw drop all forwarded traffic and jobs with open egress could not reach the internet otherwise. With Docker installed they go at the head of the `DOCKER-USER` chain; without Docker they go at the head of `FORWARD`, and only when its policy is `DROP`. Your rules for every other interface are untouched, and `inet nodus` still blocks private, link-local and metadata addresses for every job. `nodusd` puts the two rules back if they go missing and removes them when it stops. A node reports its GPUs, CPU, memory and disk, its utilization and a heartbeat. A node with no heartbeat for 10 minutes shows **Offline** and the pool’s members get one email about it. To stop new work on a host without stopping what runs there, drain it; undrain it to take work again. Deleting a node drains it, then revokes its credential: ```sh nodus node drain host-1 nodus node undrain host-1 nodus delete node/host-1 ``` ## Route jobs to the pool [Section titled “Route jobs to the pool”](#route-jobs-to-the-pool) Name the pool in a job’s placement. With `mode: Prefer` the pool is used first and the market takes what the pool cannot; with `mode: Only` work never leaves the pool. In the console, open your pool and choose **Run a job**. The job form selects your private pool under **Where to run**. Review the image, command and resources before launching. Your pool remains private to your organization. Private execution currently supports single-node Jobs, including indexed jobs. Each attempt reserves one whole host, so another attempt waits until that host is released. A GPU attempt sees only the GPUs it requested and pays the routing fee for those GPUs. Distributed jobs and other resource kinds cannot use this route yet. GPU jobs wait when the requested devices have existing GPU work, resident GPU memory, or no fresh utilization reading. After Nodus assigns a device, keep other host processes off it until the job ends; the host owner still controls processes started outside Nodus. ```yaml apiVersion: nodus.dev/v1 kind: Job metadata: name: train spec: image: nodus/pytorch:2.8-cuda12.8 command: [python, train.py] resources: {gpu: "H100:1"} placement: {pool: lab} ``` The pool’s `routing` decides what happens when it is full: | Field | Values | Meaning | | --------------- | ----------------------------- | ----------------------------------------------------------------------------------------------------- | | `waitPolicy` | `Never`, `Cheaper`, `Timeout` | Go to the market at once, wait while waiting costs less than the market, or wait up to `waitTimeout`. | | `burstToMarket` | `Allow`, `Approve`, `Deny` | Burst at market rates, wait for an approval, or never burst. | With `Approve`, a job that would burst shows `BurstApprovalRequired` until someone approves it: ```sh nodus request approve-burst job/train ``` A job on the pool is billed only the routing fee for its device-hours. A burst to the market is billed like any other job. ## Forecast demand [Section titled “Forecast demand”](#forecast-demand) Pool Predict is currently unavailable while the updated host telemetry is being qualified. Enabling it is rejected before charging. Cloud-spend projections in BYOCompute are available separately. Turn on Predict on the pool’s **Forecast** tab or with `spec.predict.enabled: true`. Forecasts start once the pool has 7 days of utilization history, then refresh hourly; recommendations refresh daily. ```sh nodus pool utilization lab nodus pool forecast lab --horizon 7d nodus pool recommendations lab ``` The forecast gives demand and free capacity in GPUs per hour with p50 and p90 bands. Recommendations cover right-sizing, idle hosts, maintenance windows, fragmentation and wait-policy tuning. Dismissing one records your reason, and the same recommendation is not raised again. Turning Predict off stops the charge at once. Host utilization history requires an updated `nodusd` reporting GPU readings. Missing samples and long gaps do not establish idle capacity, and upgrading does not reconstruct past readings. The BYOCompute **Forecast** section also shows cloud-spend projections; those use your cloud billing history and are separate from Pool Predict capacity forecasts. ## Automate actions [Section titled “Automate actions”](#automate-actions) `spec.act` sets the mode and the policies. Every action is a `PoolAction` you can list, and every executed action is audited. | Mode | What happens | | --------- | ---------------------------------------------------------- | | `Off` | Nothing is proposed. | | `Shadow` | Records what would run and runs nothing. | | `Propose` | Each action waits for approval and expires after 24 hours. | | `Auto` | Runs actions within the policies at once. | ```sh nodus get poolactions nodus pool approve lab-idle-1 nodus pool reject lab-idle-1 nodus pool revert lab-idle-1 ``` Actions only drain and undrain hosts in the pool, adjust its routing within the policy’s bounds, or approve a burst. They never stop or delete your jobs. To stop all actions at once, pause them; the pool shows `ActPaused` until you resume: ```sh nodus pool pause lab nodus pool resume lab ``` ## See your cloud accounts [Section titled “See your cloud accounts”](#see-your-cloud-accounts) To list GPU instances and spend in your AWS or GCP accounts and enroll them into pools, connect them read-only: see [Connect your cloud accounts](/docs/guides/pools/cloud-accounts/). # Connect your cloud accounts > Give Nodus read-only access to AWS or GCP to list GPU instances and spend, enroll instances into pools and revoke access at any time. A cloud account gives Nodus read-only access to your AWS account or GCP projects. Nodus lists your instances, their GPUs and your monthly spend, and you can enroll an instance into a [pool](/docs/guides/pools/). Nodus never creates, changes or deletes anything in your cloud account. Connecting a cloud account is free. ## Connect AWS [Section titled “Connect AWS”](#connect-aws) AWS access is a role in your account that only Nodus can assume, and only with an external id unique to your cloud account. 1. In the console, open **BYOCompute › Spend › Manage cloud accounts** and choose **Connect AWS**. Enter your 12-digit account ID and choose **Continue to AWS**. Nodus names the connection and prepares the role for you. Or create it from the CLI: ```sh nodus create cloudaccount aws-main --provider aws --account-id 123456789012 ``` 2. In the AWS tab, review the prefilled CloudFormation template, acknowledge the IAM permissions and choose **Create stack**. If the tab did not open, choose **Open AWS approval** in Nodus. You can also read the template with `nodus get cloudaccount/aws-main --subresource onboarding`. The stack creates a role whose policy allows only these calls: | Call | Used for | | --------------------------------------------------------------------------- | ------------------------ | | `ec2:DescribeRegions`, `ec2:DescribeInstances`, `ec2:DescribeInstanceTypes` | Instances and their GPUs | | `ce:GetCostAndUsage` | Daily spend | 3. Return to Nodus after the stack finishes. Nodus saves the role ARN, checks access automatically and opens your account when it has read your instances. There is nothing to copy. A failed check shows the reason while AWS finishes applying the permissions. Checks pause after ten minutes; **Check again** resumes the saved connection. Retrying setup reuses the existing account and its trust identity. ## Connect GCP [Section titled “Connect GCP”](#connect-gcp) GCP access is an OAuth grant limited to read-only scopes: `compute.readonly`, `bigquery.readonly` (for a billing export) and `cloud-platform.read-only`. 1. Choose **Connect GCP**, then **Continue with Google**. No account name, project ID or service account key is required. If you already know the IDs, you can enter them under **Advanced** instead. 2. Approve read-only access on Google’s page. Nodus lists the projects you can access. Choose the projects to connect, then choose **Connect projects**. Only those projects are imported. The person who signs in must complete this selection within 15 minutes. If it expires, sign in again to resume the saved connection. Nodus refuses grants with broader scopes and encrypts the stored authorization. 3. Inventory is imported independently of billing. For spend reports, enable a Cloud Billing export in the first configured project’s **nodus_billing** BigQuery dataset. The consenting Google user needs **BigQuery Job User** on that project and **BigQuery Data Viewer** on the export. Nodus detects the export table; connecting does not create an export or backfill billing history that Google has not supplied. Azure is coming soon. ## See inventory and spend [Section titled “See inventory and spend”](#see-inventory-and-spend) ```sh nodus get cloudaccounts nodus cloud inventory aws-main ``` Nodus syncs inventory hourly and spend daily. Open an account to see its observed month-to-date charges, service breakdown and daily amounts. The first spend sync reads the previous 35 days plus today’s partial day, so connecting in the middle of a month includes the earlier charges. Currencies stay separate; Nodus does not convert them. Missing days remain unknown. The date of the last successful observation stays visible when a refresh fails. AWS amounts use Cost Explorer’s unblended cost. GCP amounts use cost plus credits from the billing export, filtered to the connected projects. Charges outside those projects, including unassigned charges, are excluded. Neither is a final invoice. You continue paying your cloud directly; these amounts are separate from your Nodus balance. ### Forecast your spend [Section titled “Forecast your spend”](#forecast-your-spend) The next-30-day estimate extends the recent daily average across at least seven consecutive usable billing days. Missing days and negative billing adjustments prevent a projection. AWS non-estimated observations are preferred; provisional AWS and GCP projections exclude today and yesterday to allow for reporting lag. Each projection shows its source, history length and last included date. It assumes the recent spending pattern continues. The compute-service subtotal includes CPU and related service charges. It is not a GPU-only bill or a measured cost for an individual workload. A separate compute projection appears only when the source supports it. These cloud-bill projections are distinct from a pool’s paid Predict capacity forecasts. ### Compare a workload [Section titled “Compare a workload”](#compare-a-workload) Choose **Compare a workload with Nodus**, select matching capacity, and enter your current cost for the same workload after discounts. Enter its expected runtime on Nodus and any additional transfer, storage or ongoing commitment costs. Nodus requests a Job estimate that includes startup and shutdown; it does not start compute. The difference can be a saving or an additional cost. Changing inputs or an expired quote requires a fresh estimate. This comparison does not assume that your entire cloud bill can move to Nodus. Check that the hardware, runtime and work performed are comparable before making a migration decision. ### Run on an observed instance [Section titled “Run on an observed instance”](#run-on-an-observed-instance) The **Inventory** section lists each instance and whether it is enrolled. Choose **Enroll instance**, select a private pool in your current project, and create an install command. Run it as root on that instance. The command contains a single-use token valid for 24 hours and associates the resulting node with the observed cloud account and instance. See the [host prerequisites](/docs/guides/pools/#enroll-a-host) before installing. The installer checks the container runtime and GPU requirements before using the token. Connecting a cloud account grants observation only. Installing the agent grants execution on that host for your organization. Neither action offers capacity to other Nodus customers. ## Revoke access [Section titled “Revoke access”](#revoke-access) Disconnecting stops observation at once and deletes the credential Nodus stored: ```sh nodus delete cloudaccount/aws-main ``` Then revoke it on your side too: * **AWS**: delete the CloudFormation stack, which deletes the role. * **GCP**: remove Nodus from your Google account’s third-party access. If you revoke access on your side first, the next sync marks the account `Synced=False` with `GrantRevoked`, and the account’s members get one email about it. # Python SDK > Install the nodus-compute package, sign in, and run your first Function, Sandbox and Job from Python. The `nodus-compute` package (import `nodus`) runs Python functions, Sandboxes, Jobs and Workspaces on Nodus, and calls models through the inference data plane. It supports Python 3.10 to 3.13. ## Install and sign in [Section titled “Install and sign in”](#install-and-sign-in) ```console pip install nodus-compute nodus login ``` `nodus login` opens the console in your browser and stores a key for each org you pick in `~/.nodus/config`, which the SDK reads. On a server or in CI, create an API key in the console and set it instead: ```console export NODUS_API_KEY=nodus_sk_live_... export NODUS_PROJECT=default # optional; the project to use ``` The SDK resolves credentials in this order: arguments to `nodus.Client(...)`, the variables `NODUS_API_KEY`, `NODUS_API_URL`, `NODUS_ORG`, `NODUS_PROJECT`, `NODUS_CONTEXT` and `NODUS_CONFIG`, then the current context of `~/.nodus/config`. ## Your first Function [Section titled “Your first Function”](#your-first-function) examples/python/quickstart/app.py ```python """Quickstart: one Function called three ways. Run it with `nodus run examples/python/quickstart/app.py --n 10`. """ import nodus app = nodus.App("quickstart") @app.function(cpu=1, memory="1Gi", max_cost=1) def square(x: int) -> int: return x * x @app.local_entrypoint() def main(n: int = 10) -> None: print("remote:", square.remote(7)) # one call; blocks for the result call = square.spawn(8) # start without waiting print("spawned:", call.get(timeout=600)) print("map:", list(square.map(range(n)))) # one call per input, results in input order ``` ```console nodus run examples/python/quickstart/app.py --n 10 ``` `nodus run` creates an ephemeral App, runs `main` on your machine and deletes the App when `main` returns. Each `.remote()` call runs on Nodus and returns its result. `nodus deploy` keeps the App so other programs can call it with `nodus.Function.from_name("quickstart", "square")`. ## Blocking and asyncio [Section titled “Blocking and asyncio”](#blocking-and-asyncio) Every call blocks by default. Each one also has an `.aio` form for asyncio code: ```python async def main(): async with app.run.aio(): print(await square.remote.aio(7)) async for y in square.map.aio(range(10)): print(y) ``` ## Errors [Section titled “Errors”](#errors) Every error is a `nodus.errors.NodusError`. API errors have one class per code (`NotFound`, `Invalid`, `InsufficientCredits`, `QuotaExceeded`, …) and carry `message`, `fix`, `docs` and `request_id`. An exception raised inside a Function is raised again on your side with its own type when that type can be imported, chained from a `nodus.errors.RemoteError` that holds the remote traceback. ```python from nodus import errors try: train.remote(3e-4) except errors.InsufficientCredits as e: print(e.needed_usd, e.available_usd, e.fix) ``` ## Guides [Section titled “Guides”](#guides) * [Functions and classes](/docs/guides/python/functions/): `@app.function`, `.remote`, `.map`, `@app.cls`, deploys * [Sandboxes](/docs/guides/python/sandboxes/): exec, files, tunnels, snapshots * [Jobs](/docs/guides/python/jobs/): batch containers, outputs, multi-node gangs * [Images, Volumes and Secrets](/docs/guides/python/storage/) * [Workspaces](/docs/guides/python/workspaces/): development machines * [Agents](/docs/guides/python/agents/): durable runs, fan-out and groups * [Inference](/docs/guides/python/inference/): the OpenAI and Anthropic clients * [The resource API](/docs/guides/python/api/): `nodus.api` for any kind # Agents > Define an agent on Claude from Python, submit runs, run many in parallel, and read each run's answer and steps. An Agent is a definition: a system prompt, the Claude access it runs on and a cost cap. Each run is one conversation in its own sandbox. The [Agents guide](/docs/guides/agents/) covers the model, the cost cap and keeping a run open for follow-ups; this page is the Python side. ```python import nodus agent = nodus.ClaudeAgent("helper", system="You are careful.", families=["haiku", "sonnet"]) print(agent.remote("Use the shell to print the Python version")) # runs to the end, returns the answer run = agent.submit("Summarize the logs", keep_alive=True) run.send("message", "Now the staging logs") print(run.answer()) print(run.steps()) run.cancel() ``` `ClaudeAgent` creates the agent on first use. Pass `api_key_secret="anthropic-key"` to use your own Anthropic key, or `ClaudeAgent.from_name("claude-assistant", project="nodus")` to run the ready-made template. ## Run agents in parallel [Section titled “Run agents in parallel”](#run-agents-in-parallel) An AgentGroup runs many runs of one agent at once, with a limit on how many run together, a cost cap for the whole group and dependencies between runs. Parallel agents are in Beta, like the rest of Agents. The [Agents guide](/docs/guides/agents/#run-agents-in-parallel) explains each limit and the YAML and CLI side. examples/agents/parallel-agents/main.py ```python """Run questions in parallel under one AgentGroup, then fan a list out with agent.map. Run it with `python examples/agents/parallel-agents/main.py`. """ import nodus def main() -> None: agent = nodus.ClaudeAgent( "parallel-agents-py", system="You answer every question in one short sentence.", families=["haiku"], max_tokens=1024, per_run_max_cost="0.10", ) agent.deploy() group = nodus.AgentGroup.create("parallel-agents-py", agent, max_active=2, max_cost="0.60") try: group.submit_many( [ {"key": "a", "input": "What is the capital of France?"}, {"key": "b", "input": "What is the capital of Japan?"}, # c starts only after a and b have succeeded. {"key": "c", "input": "Say that both questions are answered.", "depends_on": ["a", "b"]}, ] ) group.seal() status = group.wait() print(f"{status.phase}: {status.counts.succeeded} of {status.counts.total} runs succeeded") for run in group.runs(): print(run.name, "->", run.answer()) finally: group.delete() # For a plain list of inputs, map does the same in one call and returns the answers in order. for answer in agent.map(["What is 2 + 2?", "What is 3 + 3?"], max_active=2): print(answer) if __name__ == "__main__": main() ``` * `nodus.AgentGroup.create(name, agent, max_active=, max_pending=, max_cost=)` creates the group over a deployed agent. `max_cost` caps the runs’ Claude usage together; their sandboxes are billed on their own. `submit_many` takes tasks `{key, input, depends_on}`, puts every task after the tasks it depends on, sends them in batches of up to 100 runs and returns the runs in your order. A cycle or a repeated key raises `nodus.errors.Invalid` before anything is sent. * `group.seal()` says that no more runs are coming. A group finishes only once it is sealed, and `group.wait()` returns its status when it has. `group.runs()` returns the member runs, and each run’s `answer()` is its text. * `group.cancel()` cancels the runs that have not finished. `group.delete()` removes the group and its runs. * `agent.map(inputs, max_active=, order_outputs=True, return_exceptions=False)` creates a group, runs one task for each input, yields each answer text and deletes the group at the end. A run that does not succeed raises `nodus.errors.AgentRunFailed`, or is yielded as the exception when you pass `return_exceptions=True`. ## Not available yet [Section titled “Not available yet”](#not-available-yet) `nodus.Agent` is the definition of a durable Python program (`@agent.entrypoint`, `@agent.step`, `ctx.step`). Runs execute on Claude, not as your code, so a definition that names `source`, an entrypoint, `setup`, `secrets`, `env`, `network`, `models`, `cpu`, `memory`, `max_cost` or worker settings raises `nodus.errors.Unsupported` when it is deployed. Submitting a run with `session_key=` or `group=` (use `AgentGroup.submit_many` for group runs), `AgentGroup.create(max_held=, evaluation=)`, `group.results()`, and a run’s `outputs`, `resolve()`, `retry()`, `suspend()`, `resume()` and `children()` raise it too. `nodus.Agent("name", image=..., per_run_max_cost=...)` and `agent.submit(input, deadline=...)` work as written, and `run.result()` and `agent.remote(input)` return the run’s answer text, the same as `run.answer()`. # The resource API > Read, watch, apply and patch any Nodus kind from Python with nodus.api, the generic client under the SDK. `nodus.api` is the generic client the rest of the SDK is built on. It works with every kind, including ones the Modal-shaped layer does not wrap, and takes the same objects as `nodus apply` and kubectl. ```python import nodus job = nodus.api.get("Job", "finetune-llama") for obj in nodus.api.list("Job", label_selector={"team": "nlp"}): print(obj["metadata"]["name"], obj["status"]["phase"]) for event in nodus.api.watch("Job", field_selector={"metadata.name": "finetune-llama"}): print(event["type"], event["object"]["status"].get("phase")) nodus.api.apply("job.yaml") # three-way merge, like nodus apply nodus.api.patch("Job", "finetune-llama", {"spec": {"maxCostUSD": "80.00"}}) estimate, etag = nodus.api.estimate(obj) # dry-run nodus.api.create(obj, if_match=etag) # launch exactly what was estimated for line in nodus.api.logs("Job", "finetune-llama", follow=True): print(line, end="") nodus.api.delete("Job", "finetune-llama") ``` Objects are plain dicts shaped like the kinds in the API reference. ## Health [Section titled “Health”](#health) ```python nodus.api.healthz() # True when the API process answers ready = nodus.api.readyz() print(ready.ready, ready.release, ready.degraded, ready.failing) ``` `readyz()` returns what `/readyz` reports instead of raising when the server is not ready: `release` is the commit it runs, `degraded` lists checks that fail without making it unready, and `failing` names the check that did. Neither call retries, so a probe sees one answer. ## Clients and retries [Section titled “Clients and retries”](#clients-and-retries) ```python client = nodus.Client(context="acme", project="nlp") # or api_key=, api_url=, org= client.get("Sandbox", "agent-1") nodus.api.set_default_client(client) # what module-level calls and the rest of the SDK use ``` Reads and watches are retried on connection errors, rate limits and `503` responses, with exponential backoff that honours `Retry-After`, for up to 5 minutes. Every create sends an `Idempotency-Key`, so a retried create never makes a duplicate. When the outcome of a write is unknown, the raised error carries `idempotency_key`: retry the same call with that key to finish it safely. # Functions and classes > Define Functions with @app.function, call them with .remote, .map and .spawn, keep state in @app.cls classes, and deploy Apps. A Function is a Python function that runs on Nodus workers. Workers start when calls arrive, stay warm for `scaledown_window` and scale between `min_workers` and `max_workers`. Every call is a FunctionCall object, so you can look it up, wait for it later or cancel it. ## Define a Function [Section titled “Define a Function”](#define-a-function) ```python import nodus app = nodus.App("finetune") image = nodus.Image.debian_slim().pip_install("torch", "transformers==4.57.6") @app.function(gpu="H100", image=image, timeout="6h", checkpoint="/nodus/state") def train(lr: float) -> dict: ... ``` | Argument | What it sets | | ---------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------- | | `gpu` | `"H100"`, `"H100:2"`, `"A100-80GB"`, `"H100!"` (exact variant), a list of alternatives, or `nodus.GPU(...)` | | `cpu`, `memory`, `ephemeral_disk` | Floors per worker; bare numbers are vCPUs and MiB | | `image` | A `nodus.Image`; defaults to `nodus/python` at your Python version | | `secrets`, `volumes`, `env`, `network` | `[nodus.Secret]`, `{"/path": nodus.Volume}`, a dict, `nodus.Egress.deny()`/`.allow(...)`/`.open()` | | `timeout`, `retries` | Per-call limit (`"6h"` or seconds); retries for exceptions (`int` or `nodus.Retries(...)`) | | `checkpoint` | A path (or `True` for `/nodus/state`) saved and restored when a worker is lost | | `interruptible`, `region`, `profile` | Allow interruptible capacity; region classes such as `["us", "eu"]`; `Balanced`, `Cost` or `Speed` | | `max_cost` | Not available for Functions yet: deploying with it is refused (Jobs and Sandboxes take it) | | `min_workers`, `max_workers`, `scaledown_window`, `target_concurrency` | Worker pool sizing (`min_containers` and `max_containers` also work) | | `name` | The member name; the Function object is `-` | Decorating does not contact Nodus, so importing the file has no side effects. The App’s code (the directory of the file, minus what `.gitignore` and `.nodusignore` exclude) is uploaded once per content hash when the App runs. ## Call it [Section titled “Call it”](#call-it) ```python with app.run(): # or: nodus run app.py result = train.remote(3e-4) # one call; blocks and returns the result call = train.spawn(1e-4) # starts a call and returns a handle print(call.get(timeout=3600)) for y in train.map([1e-4, 3e-4]): # one call per input, results in input order print(y) print(train.estimate(3e-4)) # expected cost, start time and hold, without running ``` * `.map(*iterables, order_outputs=True, return_exceptions=False)` creates calls in batches of 1,000. With `return_exceptions=True` a failed input yields its exception instead of stopping the loop. * `.starmap(pairs)` spreads each tuple into arguments; `.for_each(xs)` runs and discards the results. * `nodus.FunctionCall.from_name(name)` finds a spawned call again, from any process. * Arguments and results are serialized with cloudpickle. Values above 64 KiB travel as uploaded blobs. An exception raised in the Function is raised again in your process as its own type when it can be imported there; otherwise you get `nodus.errors.RemoteError`. Either way the remote traceback is attached. ## Deploy and look up [Section titled “Deploy and look up”](#deploy-and-look-up) ```console nodus deploy app.py ``` ```python train = nodus.Function.from_name("finetune", "train") print(train.remote(3e-4)) ``` A deploy updates the App in place: changed Functions roll their workers after in-flight calls finish, and Functions you removed from the file are deleted. ## Classes [Section titled “Classes”](#classes) `@app.cls` turns a class into one Function. `@nodus.enter()` methods run once per worker, before its first call, so expensive setup such as loading a model happens once; `@nodus.exit()` methods run when the worker drains. Methods marked `@nodus.method()` get `.remote()`, `.map()` and `.spawn()`. examples/python/classes/app.py ```python """A class whose model loads once per worker, then serves many calls. Run it with `nodus run examples/python/classes/app.py`. """ import nodus app = nodus.App("classes") @app.cls(cpu=2, memory="4Gi", scaledown_window="5m", max_cost=1) class Greeter: @nodus.enter() def load(self) -> None: # Runs once when a worker starts, before its first call: load weights or open connections here. self.greeting = "hello" @nodus.method() def greet(self, name: str) -> str: return f"{self.greeting}, {name}" @nodus.exit() def close(self) -> None: self.greeting = "" @app.local_entrypoint() def main() -> None: greeter = Greeter() print(greeter.greet.remote("Ada")) print(list(greeter.greet.map(["Grace", "Linus"]))) ``` Classes take no constructor arguments; configure them in the `enter` hook. ## Clustered Functions (Beta) [Section titled “Clustered Functions (Beta)”](#clustered-functions-beta) Beta Clustered Functions need the distributed training beta for your org. Checkpointed resume of a gang is not yet qualified. `@nodus.clustered(size=N)` below `@app.function` runs each `.remote()` or `.spawn()` as a gang of `N` nodes. Every member runs the function; rank 0’s return value is the result. `nodus.cluster.info()` tells each member its rank, the gang size, the member addresses and the rendezvous address, so `torch.distributed` initializes from the environment. examples/functions/clustered/app.py ```python """A clustered Function: each call runs as one gang of two nodes, and rank 0's return value is the result (Beta). Run it with `nodus run examples/functions/clustered/app.py`. """ import nodus app = nodus.App("clustered") image = nodus.Image.from_registry("nodus/pytorch:2.8-cuda12.8") @app.function(gpu="H100", image=image, timeout="30m", max_cost=1) @nodus.clustered(size=2) def whoami() -> dict: info = nodus.cluster.info() # rank, size, member addresses, rendezvous address and epoch print(f"rank {info.rank} of {info.size}; rendezvous {info.master_addr}:{info.master_port}") return {"rank": info.rank, "size": info.size} @app.local_entrypoint() def main() -> None: result = whoami.remote() assert result == {"rank": 0, "size": 2}, result print(result) ``` `network` (`Colocated`, `Regional`, `Global`) and `transport` (`Direct`, `Auto`) choose where members may be placed. Clustered Functions keep no warm workers, and `.map()` on them raises `nodus.errors.Unsupported`. # Inference > Call catalog models with the official OpenAI and Anthropic clients through nodus.llm, and create named InferenceEndpoints. The inference data plane speaks the OpenAI and Anthropic wire formats. `nodus.llm` returns the official clients, configured with your key and the Nodus base URL, so every feature of those SDKs works unchanged. ```console pip install "nodus-compute[openai]" # or [anthropic] ``` examples/python/inference/chat.py ```python """Chat with a catalog model through the OpenAI-compatible inference data plane. Needs `pip install "nodus-compute[openai]"`. Run it with `python examples/python/inference/chat.py`. """ import nodus def main() -> None: client = nodus.llm.openai() # the official OpenAI client with your Nodus key and base URL reply = client.chat.completions.create( model="nodus/gpt-oss-20b", messages=[{"role": "user", "content": "Say hello in five words."}], max_tokens=32, ) print(reply.choices[0].message.content) if __name__ == "__main__": main() ``` ```python client = nodus.llm.openai(project="nlp") # usage is attributed to the project claude = nodus.llm.anthropic() # messages API aclient = nodus.llm.async_openai() # AsyncOpenAI ``` Inside a Job, Function or Sandbox, `nodus.llm` uses the container’s built-in proxy: no key is needed, and the calls are billed to, and capped by, the run that makes them. ## Named endpoints [Section titled “Named endpoints”](#named-endpoints) An InferenceEndpoint gives a model its own base URL with rate limits, allowed keys and a spending cap. ```python ep = nodus.InferenceEndpoint.create("support-bot", "nodus/gpt-oss-120b", rpm=120, max_concurrent=8, max_cost=50) client = ep.openai() print(ep.usage()) # requests, tokens and cost over the last 24 hours ``` # Jobs > Run a container to completion from Python with your code uploaded, follow its logs, download its outputs and run multi-node gangs. A Job runs a command to completion. Nodus picks the capacity, keeps its checkpoints and recovers it if the capacity is reclaimed; `max_cost` caps what it may spend. examples/python/jobs/main.py ```python """Upload this directory as a Job's source, follow its logs and download its output. Run it with `python examples/python/jobs/main.py`. """ from pathlib import Path import nodus HERE = Path(__file__).parent def main() -> None: job = nodus.Job.run( image="nodus/python:3.12", command=["python", "train.py"], source=HERE, # uploaded once as a content-addressed blob; .gitignore and .nodusignore apply cpu=2, memory="4Gi", timeout="30m", max_cost=1, outputs={"report": "/nodus/outputs/report.txt"}, ) print("estimate:", job.estimate()) for line in job.logs(follow=True): print(line, end="") job.wait() # raises nodus.errors.JobFailed with the exit code and log tail print("saved", job.outputs["report"].download(HERE / "report.txt")) if __name__ == "__main__": main() ``` ## Create [Section titled “Create”](#create) ```python job = nodus.Job.run( name="finetune-llama", # optional; a generated name otherwise image="nodus/pytorch:2.8-cuda12.8", command=["python", "train.py", "--epochs", "3"], source=".", # uploads the directory; or {"repo": "acme/trainer", "ref": "main"} gpu="H100", secrets=["hf-token"], max_cost=40, timeout="12h", expected_duration="6h", checkpoint="/nodus/state", interruptible=True, region=["us", "eu"], outputs={"adapter": "/nodus/outputs/adapter"}, ) ``` The source upload skips what `.gitignore` and `.nodusignore` exclude and refuses more than 500 MiB compressed unless you pass `allow_large_source=True`. Everything written under `/nodus/outputs` is collected when the Job succeeds. ## Watch and collect [Section titled “Watch and collect”](#watch-and-collect) ```python print(job.estimate()) # expected cost p50 and p90, start time and the hold for line in job.logs(follow=True): print(line, end="") job.wait() # raises nodus.errors.JobFailed(job, exit_code, log_tail) job.outputs["adapter"].download("./adapter") # checks the sha256, then renames into place ``` `job.suspend()`, `job.resume()` and `job.cancel()` change the Job’s state; `job.attempts()` lists every attempt with its placement; `job.exec("nvidia-smi")` runs a command in the running container. ## Multi-node gangs (Beta) [Section titled “Multi-node gangs (Beta)”](#multi-node-gangs-beta) Beta Gangs need the distributed training beta for your org. Checkpointed resume of a gang is not yet qualified. ```python ddp = nodus.Job.run( name="ddp-smoke", image="nodus/pytorch:2.8-cuda12.8", command=["torchrun", "train.py"], gpu="H100", max_cost=10, distributed=nodus.Distributed(nodes=2, launcher="Torchrun", network="Global", transport="Auto"), ) for line in ddp.logs(follow=True, rank="all"): # every rank, prefixed [r0], [r1] print(line, end="") ``` Every rank gets the rendezvous environment (`MASTER_ADDR`, `WORLD_SIZE`, torchrun’s `PET_*` variables), and `nodus.cluster.info()` reads it for you. `nodus.checkpoint.dcp.save(state, step)` and `.load(state)` write and read `torch.distributed.checkpoint` shards to the gang’s checkpoint storage. # Sandboxes > Create isolated containers from Python, run commands in them, move files and snapshot their filesystem. A Sandbox is an isolated, long-running container for agents and untrusted code. You drive it with commands and file operations; it stops when idle and starts again on the next command. examples/python/sandbox/main.py ```python """Create a Sandbox, run commands in it, move files and delete it. Run it with `python examples/python/sandbox/main.py`. """ import nodus def main() -> None: sb = nodus.Sandbox.create( cpu=1, memory="2Gi", idle_timeout="5m", max_cost=1, ) try: p = sb.exec("python", "-c", "print(6 * 7)") print("stdout:", p.stdout.read().strip(), "exit:", p.wait()) with sb.open("/workspace/notes.txt", "w") as f: f.write("written from the SDK\n") print(sb.exec("cat", "/workspace/notes.txt").stdout.read(), end="") finally: sb.terminate() if __name__ == "__main__": main() ``` ## Create [Section titled “Create”](#create) ```python sb = nodus.Sandbox.create( name="agent-1", # optional; creating the same name again reconnects image="nodus/agent-tools", # the default: Python 3.12, Node 22, git cpu=2, memory="4Gi", timeout="24h", idle_timeout="5m", on_idle="stop", network=nodus.Egress.open(), # outbound access; the default is nodus.Egress.deny() max_cost=5, ) ``` `create` returns as soon as the Sandbox is admitted; commands wait for it to start. Egress is denied unless you open it. `nodus.Sandbox.from_name("agent-1")` reconnects and starts it again if you stopped it. Not available yet These arguments raise `nodus.errors.Unsupported` before anything is sent: `volumes=`, `ports=` (and `sb.tunnels.open()`), `init=`, `service=`, `gpu=` and an allow-list egress (`nodus.Egress.allow(...)`). `image=` takes a published image name, not an `nodus.Image` that Nodus builds. `secrets=[...]` is sent when you set it, and raises `Unsupported` while the API does not accept secrets on a Sandbox. ## Run commands [Section titled “Run commands”](#run-commands) ```python p = sb.exec("python", "-c", "print(6 * 7)") print(p.stdout.read()) # everything the process wrote to stdout assert p.wait() == 0 # the exit code p = sb.exec("pip install requests && python app.py") # one string runs under /bin/sh -c for chunk in p.stdout: # stream output as it arrives print(chunk, end="") p = sb.exec("bash", pty=True) p.write("ls\n"); p.resize(40, 120); p.signal("SIGINT") ``` Every command is recorded as a Process, and its output is kept, so a reader that reconnects continues where it stopped. `p.stdin.write(...)` and `p.stdin.write_eof()` feed input; `p.cancel()` stops the process. ## Files [Section titled “Files”](#files) ```python with sb.open("/workspace/notes.txt", "w") as f: f.write("hi") print(sb.files.read("/workspace/notes.txt")) print(sb.files.list("/workspace")) ``` ## Stop, start and snapshot [Section titled “Stop, start and snapshot”](#stop-start-and-snapshot) ```python sb.stop() # saves the filesystem; compute billing stops, the saved files count as storage sb.start() image = sb.snapshot_filesystem() # Beta: an Image of the current filesystem sb.terminate() # deletes it ``` Beta `snapshot_filesystem()` is Beta. Memory is not captured: processes start fresh in a Sandbox created from the image. A Sandbox cannot start from the snapshot Image yet. # Images, Volumes and Secrets > Build container images from Python, share files between workers with Volumes, and pass credentials with Secrets. ## Images [Section titled “Images”](#images) ```python image = ( nodus.Image.from_registry("nodus/pytorch:2.8-cuda12.8") .apt_install("git") .pip_install("transformers==4.57.6", "peft") .env({"HF_HUB_ENABLE_HF_TRANSFER": "1"}) .add_local_python_source("mylib") ) image = nodus.Image.from_dockerfile("Dockerfile", context=".") image = nodus.Image.debian_slim("3.12").uv_pip_install("numpy") ``` Nothing is built until an App that uses the image runs, or until you call `image.build()`, which streams the build log. An Image is named by the hash of its steps, so the same chain reuses the same build. `add_local_file`, `add_local_dir` and `add_local_python_source` upload local files into the image; `uv_sync(".")` installs a uv project from its lockfile. ## Volumes [Section titled “Volumes”](#volumes) ```python vol = nodus.Volume.from_name("data", create_if_missing=True) vol.put_file("./local.csv", "/train/local.csv") print(vol.listdir("/train")) with vol.batch_upload() as up: up.put_directory("./dataset", "/train") weights = nodus.Volume.import_from("llama", huggingface="meta-llama/Llama-3.1-8B", revision="0e9e39f") ``` A Volume made by `create_if_missing=True` is `ReadWriteMany`, with Modal’s semantics: each worker mounts the latest revision when it starts, `vol.commit()` publishes that worker’s changed files as a new revision (the last writer wins per path), and `vol.reload()` mounts the newest revision. Close open files before `reload()`. Uploads and downloads use the `nodus` CLI that the package installs. examples/python/storage/app.py ```python """Workers share a ReadWriteMany Volume: each commits its result, and a reload sees everyone's. Run it with `nodus run examples/python/storage/app.py`. """ from pathlib import Path import nodus app = nodus.App("storage") results = nodus.Volume.from_name("storage-example", create_if_missing=True) token = nodus.Secret.from_dict({"GREETING": "hello"}) @app.function(volumes={"/results": results}, secrets=[token], max_workers=4, max_cost=1) def work(i: int) -> str: import os Path(f"/results/{i}.txt").write_text(f"{os.environ['GREETING']} {i}\n") results.commit() # publish this worker's files as a new revision return f"{i}.txt" @app.function(volumes={"/results": results}, max_cost=1) def collect() -> list[str]: results.reload() # mount the latest revision return sorted(p.name for p in Path("/results").iterdir()) @app.local_entrypoint() def main() -> None: print(list(work.map(range(4)))) print(collect.remote()) ``` ## Secrets [Section titled “Secrets”](#secrets) ```python hf = nodus.Secret.from_name("hf-token") # an existing Secret cfg = nodus.Secret.from_dict({"API_TOKEN": "..."}) # created with the App and deleted with it env = nodus.Secret.from_dotenv(".env") nodus.Secret.create("hf-token", {"HF_TOKEN": "hf_..."}) # writes a new version if it exists ``` Every key arrives in the container as an environment variable and as a file under `/run/secrets//`. Values are never returned by the API, and running containers keep the version they started with. # Workspaces > Create a development machine with SSH, VS Code, JupyterLab and a persistent home from Python. A Workspace is a development machine with a persistent home directory. It stops when idle, and its home is saved when it stops. ```python lab = nodus.Workspace.create("lab", gpu="H100:2", image="nodus/workspace-pytorch-cuda", idle_timeout="2h") lab.wait_ready() url = lab.open("vscode") # or "jupyter"; opens the browser and returns the URL lab.ssh() # an SSH session through the nodus CLI lab.schedule(ready_by="2026-10-01T08:45:00-07:00", stop_at="2026-10-01T19:00:00-07:00") lab.stop() for s in lab.sessions(): # one entry per running period, with what it cost print(s.billed_seconds, s.charge_usd, s.stop_reason) ``` * The home Volume is `volume=` or `-home`, created if it does not exist. `ephemeral=True` skips it: files are lost on stop. * `lab.start()` wakes it after an idle or schedule stop; an SSH connection also wakes it. * `lab.exec(...)` and `lab.files` work as on Sandboxes. # Sandboxes > Create an isolated container, run commands in it with streaming output, read and write files, and stop it to stop paying for compute. A Sandbox is an isolated, long-running container for agents and untrusted code. You drive it with commands and file requests. When nothing uses it for a while it stops, keeps its `/workspace` directory and stops costing compute; the next command or file request starts it again. You pay per second while it holds compute and nothing while it is stopped. Sandboxes run on CPU machines Nodus operates, each in its own isolated runtime with no network access unless you open it. [Sandbox isolation](/docs/concepts/sandboxes-isolation/) explains what keeps one Sandbox from another. ## Quick start [Section titled “Quick start”](#quick-start) This manifest is the whole Sandbox. It asks for 1 vCPU and 2 GiB of memory, stops after 5 idle minutes and caps its spending at 1 USD: sandbox.yaml ```yaml apiVersion: nodus.dev/v1 kind: Sandbox metadata: name: hello spec: image: nodus/agent-tools # required: Python 3.12, Node 22 and git resources: cpu: "1" memory: 2Gi network: egress: policy: Deny # the default: no outbound traffic lifecycle: idleTimeout: 5m # stop after 5 minutes with no exec, file request or running process onIdle: Stop # keep /workspace and stop paying for compute; Delete removes the Sandbox instead maxLifetime: 2h # delete it 2 hours after creation whatever it is doing maxCostUSD: "1.00" ``` Create it, run a command, copy a file in and read it back: ```console $ nodus apply -f sandbox.yaml sandbox/hello created $ nodus exec sb/hello -- echo hello from a sandbox hello from a sandbox $ nodus cp ./notes.txt sb/hello:/workspace/notes.txt $ nodus exec sb/hello -- cat /workspace/notes.txt written from my laptop $ nodus stop sb/hello $ nodus exec sb/hello -- cat /workspace/notes.txt # starts it again; /workspace is kept written from my laptop $ nodus delete sb/hello ``` From Python, the same steps are in the [Python Sandboxes guide](/docs/guides/python/sandboxes/). ## Create a Sandbox [Section titled “Create a Sandbox”](#create-a-sandbox) `nodus apply -f sandbox.yaml`, `nodus create sandbox ` and the SDK all create the same object. `image` is required on the API; the CLI and the SDK fill in `nodus/agent-tools` (Python 3.12, Node 22 and git) when you leave it out. | Field | Default | Accepted values | | ---------------------------- | ------------- | --------------------------------------------------------------------------- | | `spec.image` | — (required) | Any image reference, or a Nodus catalog image such as `nodus/agent-tools` | | `spec.resources.cpu` | `1` | 0.25 to 64 vCPU | | `spec.resources.memory` | `1Gi` | 128 MiB to 256 GiB | | `spec.resources.disk` | `10Gi` | 1 GiB to 200 GiB | | `spec.network.egress.policy` | `Deny` | `Deny` (no outbound traffic) or `Open` (public internet) | | `spec.lifecycle.idleTimeout` | `5m` | `0s` (never idle) or 1 minute to 24 hours | | `spec.lifecycle.onIdle` | `Stop` | `Stop` or `Delete` | | `spec.lifecycle.maxLifetime` | `24h` | 1 minute to 720 hours | | `spec.continuity.mode` | `Snapshotted` | `Snapshotted` (keep `/workspace` across stops) or `Ephemeral` (start empty) | | `spec.workingDir` | `/workspace` | An absolute path | | `spec.maxCostUSD` | none | A USD amount such as `"5.00"`; the Sandbox stops when it has cost this much | | `spec.state` | `Running` | `Running` or `Stopped` | `nodus get sb/hello` shows `PHASE` (`Pending`, `Starting`, `Running`, `Stopping`, `Stopped`, `Recovering`, `Terminating`, `Failed`), `ACTIVITY` (`Busy` or `Idle`), the `CPU`, `MEMORY` and `GPU` you asked for, and `COST` so far. ### Create by name reconnects [Section titled “Create by name reconnects”](#create-by-name-reconnects) Creating a Sandbox whose name already exists in the project does not make a second one: * **Same spec.** You get the existing Sandbox back, and if it is stopped it starts again. This is how an agent reconnects to its Sandbox after a restart: it creates the same manifest every time. * **A `maxCostUSD` that is added or raised.** The new budget is applied; nothing else changes. A budget can only be raised: a lower one, or none when the Sandbox has one, counts as a different spec. * **Any other change.** The request fails with `409 AlreadyExists` and lists the fields that differ in `details.diff` and `details.causes`. Delete the Sandbox and create it again to change its image, shape or lifecycle. GPU Sandboxes are not available yet: a spec with `resources.gpu` is refused. ## Run commands [Section titled “Run commands”](#run-commands) `nodus exec` runs one command in the running Sandbox and streams its standard output and standard error back as the command writes them. The exit code of `nodus exec` is the command’s. Every command is recorded as a Process of the Sandbox, whatever started it (the CLI, an agent or the MCP tools): its name, such as `sb-hello-12`, comes back with the result, and `nodus get processes` lists it with how it ended. ```console $ nodus exec sb/hello -- python -c 'print(6 * 7)' 42 $ nodus exec sb/hello -- sh -c 'pip list 2>/dev/null | head -3; exit 3'; echo "exit $?" Package Version ---------- ------- pip 24.2 exit 3 ``` A command runs in `spec.workingDir` (`/workspace` unless you set it) with the Sandbox’s environment, as a non-root user. Pass a single string to `sh -c` to use pipes and redirection. Every command counts as activity, so a Sandbox never stops as idle while a command is running. ### Open a shell [Section titled “Open a shell”](#open-a-shell) `nodus shell` is `nodus create sandbox` and `nodus exec -it` in one step: it opens an interactive terminal in a Sandbox and creates the Sandbox first when the name is new. It takes the same flags as `nodus create sandbox`, and the image is `nodus/agent-tools` unless you pass `--image`. ```console $ nodus shell sandbox/dev --cpu 2 --memory 4Gi sandbox/dev created sandboxes/dev is starting; waiting for it ``` You are now in a terminal in the Sandbox, in `/workspace`. When you exit, `nodus shell sandbox/dev` opens it again with `/workspace` as you left it. The Sandbox stays after you exit, and stops by itself when idle. Pass `--rm` to delete a Sandbox this command created when the shell exits; `--rm` refuses a Sandbox that already exists, and so does a flag such as `--cpu`, because neither changes a Sandbox that is there. `nodus shell` with no name gives the shell a Sandbox of its own and deletes it on exit unless you pass `--keep`. Put a command after `--` to run it instead of `bash`: `nodus shell sandbox/dev -- zsh`. With input that is not a terminal, the shell runs without one, so `echo 'make test' | nodus shell sandbox/dev` works in a script. The exit code is the shell’s. ### Timeouts [Section titled “Timeouts”](#timeouts) Each command has its own timeout: 10 minutes unless you set one, and at most 24 hours. When the timeout passes, the process is killed with `SIGKILL` and the command ends with the reason `DeadlineExceeded`. The Sandbox keeps running. Set the timeout per command, for example `sb.exec("make", "test", timeout="2m")` in Python or `"timeout": "2m"` in the API request. How a command ended is in its result: | Reason | Meaning | | --------------------- | ------------------------------------------------------------------ | | none | Exit code 0 | | `NonZeroExit` | The command exited with another code | | `Signaled` | A signal ended it; the exit code is 128 plus the signal number | | `DeadlineExceeded` | Its timeout passed and it was killed | | `StartFailed` | It could not start, for example because the program does not exist | | `NodeLost` | The machine running the Sandbox was lost while the command ran | | `OutputLimitExceeded` | It wrote more output than a Process may keep | | `ParentStopped` | The Sandbox stopped or was deleted while the command ran | | `Cancelled` | You cancelled it | A command line holds at most 1,024 arguments and 64 KiB in total, plus up to 256 extra environment variables. A Sandbox runs at most 1,024 commands at once; the next one fails with `429 TooManyRequests` until one ends. ## Files [Section titled “Files”](#files) `nodus cp` copies files in both directions, and works on every image because the file service is part of the Sandbox runtime, not of the image: ```console $ nodus cp ./notes.txt sb/hello:/workspace/notes.txt # laptop to Sandbox $ nodus cp sb/hello:/workspace/result.json ./result.json # Sandbox to laptop ``` The same operations are one request each on `/apis/nodus.dev/v1/namespaces/{project}/sandboxes/{name}/files`: | Request | Does | | -------------------------------------- | -------------------------------------------------------------------- | | `GET …/files?path=/workspace/a.txt` | Returns the file’s bytes | | `PUT …/files?path=/workspace/a.txt` | Replaces the file with the request body, creating parent directories | | `DELETE …/files?path=/workspace/a.txt` | Deletes a file or an empty directory | A write replaces the file atomically: a reader sees the old content or the new content, never half of it. A file is at most 64 MiB per request; a larger one fails with `413 FileTooLarge`. Paths are absolute, and a path that does not exist answers `404 NotFound`. ### Writing without overwriting someone else’s change [Section titled “Writing without overwriting someone else’s change”](#writing-without-overwriting-someone-elses-change) When two writers share a file, send the digest of the content you started from as `expectedSHA256`: ```console $ curl -X PUT --data-binary @plan.md "$API/sandboxes/hello/files?path=/workspace/plan.md&expectedSHA256=$(sha256sum old-plan.md | cut -d' ' -f1)" ``` The write happens only if the file still has that digest. If it changed, nothing is written and the request fails with `412 SHA256Mismatch`: read the file again, merge, and retry. The digest is 64 lowercase hex digits. The digest of empty content (`e3b0c442…b855`) matches a file that does not exist yet, so it means “create only”. ## Stop, start and wake [Section titled “Stop, start and wake”](#stop-start-and-wake) A Sandbox is `Running` while it holds compute and `Stopped` while it does not. A stop saves `/workspace` first, so a later start gives it back, as long as `spec.continuity.mode` is `Snapshotted`, the default. ```console $ nodus stop sb/hello # saves /workspace, releases compute, ends the charge $ nodus start sb/hello # starts it again with /workspace restored $ nodus get sb/hello NAME PHASE ACTIVITY CPU MEMORY GPU COST AGE hello Stopped Idle 1 1Gi - $0.0042 12m ``` `nodus stop` sets `spec.state: Stopped` and the Sandbox stays stopped until you start it. Every other stop is made by the platform, shows its reason in `status.stopReason`, and ends the next time something needs the Sandbox: | `stopReason` | Why it stopped | | --------------------- | --------------------------------------------- | | `User` | You ran `nodus stop` | | `Idle` | Nothing used it for `idleTimeout` | | `InsufficientCredits` | Your balance could not fund the next period | | `BudgetExceeded` | A Budget on the project reached its limit | | `MaxCostReached` | The Sandbox cost as much as `spec.maxCostUSD` | A stopped Sandbox starts again on any of these wake triggers: * a command (`nodus exec`, or an exec request), * a file request, * `nodus start`, * creating a Sandbox with the same name and an identical spec. Credits and budgets still apply to a wake. A Sandbox stopped for `InsufficientCredits`, `BudgetExceeded` or `MaxCostReached` stays stopped until the balance, the Budget or `maxCostUSD` allows the next period of compute. Until then a command, a file request and `nodus start` answer `402` instead of `SandboxStarting`: `InsufficientCredits` states the hold the start needs and what you have, and `BudgetExceeded` the Budget or `maxCostUSD` that is full. Add credits, raise the Budget or raise `maxCostUSD`, then send a trigger again. A create by name records its wake too, but returns the Sandbox as it is: still stopped, with `Funded` false in its conditions. ### `SandboxStarting` and `Retry-After` [Section titled “SandboxStarting and Retry-After”](#sandboxstarting-and-retry-after) A command or file request that finds the Sandbox stopped starts it and answers `503 SandboxStarting` with a `Retry-After` header (2 seconds), because the container is not ready yet. The CLI and the SDK wait and retry for you, so `nodus exec` on a stopped Sandbox just takes a few seconds longer. If you call the API directly, retry after the time in `Retry-After`; each retry is also a wake trigger, so it does no harm to send them early. Other reasons a request can fail while the Sandbox is not usable: | Status and reason | Meaning | | ------------------------- | ---------------------------------------------------------------------------------------- | | `409 SandboxNotRunning` | The Sandbox is being deleted | | `409 SandboxFailed` | The Sandbox failed to start or lost its machine; `status.reason` says why | | `402 InsufficientCredits` | The Sandbox is stopped for lack of credits and cannot start yet | | `402 BudgetExceeded` | The Sandbox is stopped because a Budget or its `maxCostUSD` is full and cannot start yet | | `404 NotFound` | There is no Sandbox with that name | Each code has a page in the [error reference](/docs/reference/errors/). ## Idle, `onIdle` and `maxLifetime` [Section titled “Idle, onIdle and maxLifetime”](#idle-onidle-and-maxlifetime) A Sandbox is **idle** when no command is running, no file request arrived for `idleTimeout` and no terminal had input in that time. A running process, including one that prints nothing for hours, counts as busy; a terminal that is open but silent does not count as activity. When the idle time passes, `onIdle` decides what happens: * `Stop` (the default) saves `/workspace`, releases compute and sets `stopReason: Idle`. The next wake trigger starts it again. * `Delete` removes the Sandbox with its files. Use it for one-off work that leaves nothing worth keeping. `idleTimeout: 0s` turns idle detection off, so the Sandbox runs until you stop it, until `maxLifetime` or until a budget ends it. `maxLifetime` is a hard deadline counted from creation, 24 hours by default. When it passes, the Sandbox is deleted whatever it is doing, running or stopped. It is the backstop that keeps a forgotten Sandbox from living forever. Create a new Sandbox, or one with a longer `maxLifetime`, to carry on. ## Network access [Section titled “Network access”](#network-access) A Sandbox has no outbound network access unless you open it. `spec.network.egress.policy` is one of: * `Deny` (the default): nothing leaves the Sandbox, so code in it cannot download or upload anything. * `Open`: the Sandbox can reach the public internet. It still cannot reach other Sandboxes, the machine it runs on or Nodus’s own services. A Sandbox accepts no inbound connections; you reach it only through commands and file requests. ## What a Sandbox costs [Section titled “What a Sandbox costs”](#what-a-sandbox-costs) A Sandbox is billed per second, from the moment its machine slot is reserved until it is released, at the rate of its shape. The rate is the Nodus machine price for the CPU, memory and disk you asked for, and it already includes Nodus’s 12.5 % ([how pricing works](/docs/concepts/pricing/)). Four things to know: * **Starting and stopping are billed.** The seconds spent starting (or restoring `/workspace`) and the seconds spent saving a stop are part of the reservation, shown as their own lines on the bill. * **Stopped time is free.** A stopped Sandbox holds no compute, so no compute cost accrues. The saved `/workspace` counts as storage, which is metered separately from compute. * **The rate is frozen per start.** The rate a start shows is the rate that start pays, even if prices change while it runs. * **Shared capacity has limits.** New shared capacity starts for you at most four times an hour, and none starts for 30 minutes after capacity started for you goes unused; until your organization’s first purchase it starts only for Sandboxes of up to 3 vCPU and 11 GiB. Otherwise a Sandbox runs on a machine of its own, at that machine’s rate. `nodus get sb/hello` shows the running total in the `COST` column. Set `spec.maxCostUSD` to cap one Sandbox: it stops when it has cost that much, and you can raise the cap with another create by name. Without a cap, your credit balance and any Budget on the project are the limits. ## Next steps [Section titled “Next steps”](#next-steps) * [Sandbox isolation](/docs/concepts/sandboxes-isolation/) for what separates your Sandbox from the machine and from other Sandboxes. * [Sandboxes from Python](/docs/guides/python/sandboxes/) for the SDK, including terminals and previews. # Secrets > Store API tokens, credentials and registry logins once, and pass them to Jobs, Sandboxes, Functions and Agents as env vars and files. A Secret holds named values, such as an API token or a database URL, encrypted with a key that belongs to your org. You create it once and name it in the Jobs and Sandboxes that need it. Nodus never returns a value through the API, the console or the CLI, and it redacts secret values from logs and outputs. ## Create a Secret [Section titled “Create a Secret”](#create-a-secret) ```console nodus secret create hf-token --from-literal HF_TOKEN=hf_xxx nodus secret create app-env --from-env-file .env nodus secret create registry --type Registry \ --from-literal server=ghcr.io --from-literal username=ada --from-literal password=ghp_xxx ``` Key names start with a letter or `_`, use letters, digits and `_`, and may not start with `NODUS_`. A Secret holds up to 64 keys of at most 4 KiB each. `nodus get secret hf-token` shows the key names and the current version, never the values. ## Use it in a Job [Section titled “Use it in a Job”](#use-it-in-a-job) List the Secret under `secrets` to receive every key as an env var and as a read-only file `/run/secrets//`, or pick one key under another name with `valueFrom.secretKeyRef`: job.yaml ```yaml apiVersion: nodus.dev/v1 kind: Job metadata: name: ex-secrets-env spec: image: nodus/python:3.12 secrets: [ex-secrets-env] # every key as an env var and a file under /run/secrets/ex-secrets-env/ env: - name: TOKEN # one key under another name valueFrom: {secretKeyRef: {name: ex-secrets-env, key: API_TOKEN}} command: [python, check.py] ``` check.py ```python import os token = os.environ["API_TOKEN"] with open("/run/secrets/ex-secrets-env/API_TOKEN") as f: from_file = f.read() # Print facts about the value, never the value: Nodus redacts secret values from logs anyway. print(f"token has {len(token)} characters") print("file matches" if from_file == token else "file differs") print("alias matches" if os.environ["TOKEN"] == token else "alias differs") ``` ```console nodus apply -f job.yaml nodus logs job/ex-secrets-env ``` The same `secrets` and `env` fields work on Sandboxes, Workspaces, Functions and Agents. In Python: ```python hf = nodus.Secret.from_name("hf-token") env = nodus.Secret.from_dotenv(".env") ``` A command whose literal `env` value equals one of the Secret values in its project is refused with `SecretValueInEnv`: reference the Secret instead of pasting its value. The same check applies to commands you start in a running Sandbox, Workspace or Job with `nodus exec` or the Processes API, which also refuse an `env` that sets the same name twice. Values shorter than 8 characters, and the `server` and `username` of a Registry Secret, are not checked. A Job receives the Secret versions pinned when you submitted it, and a Sandbox the versions pinned when it started. If a Secret it uses is deleted (even if you create a new one with the same name), or a new value replaced a Job’s pinned version before a retry started, that attempt fails with `LaunchFailed`. Secret values delivered as env vars may total at most 1 MiB per container, and as files at most 8 MiB. ## Change a value [Section titled “Change a value”](#change-a-value) Writing a Secret again creates a new version; writing the same values again changes nothing. A Sandbox pins the newest version each time it starts, so it gets the new value on its next start, while a running Sandbox keeps the version it started with. A Job keeps the versions pinned when you submitted it, so every attempt runs with the same values; a retry that starts after you replace a value fails with `LaunchFailed`, so submit the Job again to use it. Earlier versions are deleted 7 days after they are replaced. hf-token.yaml ```yaml apiVersion: nodus.dev/v1 kind: Secret metadata: {name: hf-token} spec: stringData: {HF_TOKEN: hf_new} ``` ```console nodus apply -f hf-token.yaml ``` ## Private registries [Section titled “Private registries”](#private-registries) A Secret with `type: Registry` holds `server`, `username` and `password`. Name it under `imagePullSecrets` to run or build from a private image: ```yaml spec: image: ghcr.io/acme/trainer:1.4 imagePullSecrets: [{name: registry}] ``` Nodus resolves the tag to a digest when you submit, using the Secret, and the machine that runs the work pulls that digest with the same Secret version. The credential is used for the pull alone: it is not an env var or a file in the container. When the work runs on a provider that pulls containers itself, Nodus first copies the image by digest into your org’s space in the Nodus registry, so your registry credential never leaves Nodus. ## Limits and errors [Section titled “Limits and errors”](#limits-and-errors) | Error | Meaning | Fix | | ------------------ | ------------------------------------ | -------------------------------------------- | | `SecretValueInEnv` | An `env` value equals a Secret value | Use `valueFrom.secretKeyRef` or `secrets` | | `EncryptionFailed` | The value could not be encrypted | Retry; the Secret keeps its previous version | | `ImagePullFailed` | The registry refused the pull Secret | Check `server`, `username` and `password` | # Sweeps > Run one Job across GPU types, regions and parameter values, and compare cost, time and throughput per cell. A Sweep runs the same Job once for every combination of the values you list: GPU types, region classes and your own parameters. Each combination is a cell. When the cells finish, the Sweep shows what each one cost, how long it took and how fast it went, and points out the cheapest and the fastest. ## Submit a Sweep [Section titled “Submit a Sweep”](#submit-a-sweep) This Sweep measures training throughput for three batch sizes on two GPU types, six cells in all: sweep.yaml ```yaml apiVersion: nodus.dev/v1 kind: Sweep metadata: name: batch-size spec: maxCostUSD: "0.75" maxParallel: 2 # 2 GPU types x 3 batch sizes = 6 cells, each a Job named batch-size-. matrix: gpu: [L4, A10] params: BATCH_SIZE: ["32", "64", "128"] template: kind: Job spec: image: nodus/pytorch timeout: 10m command: - python - -c - | import os, time, torch bs = int(os.environ["NODUS_PARAM_BATCH_SIZE"]) model = torch.nn.Sequential(torch.nn.Linear(1024, 4096), torch.nn.ReLU(), torch.nn.Linear(4096, 10)).cuda() opt = torch.optim.AdamW(model.parameters()) x, y = torch.randn(bs, 1024).cuda(), torch.randint(0, 10, (bs,)).cuda() start, steps = time.time(), 300 for _ in range(steps): opt.zero_grad() torch.nn.functional.cross_entropy(model(x), y).backward() opt.step() torch.cuda.synchronize() print(f"batch {bs}: {steps * bs / (time.time() - start):.0f} samples/s") ``` ```console $ nodus apply -f sweep.yaml sweep.nodus.dev/batch-size created ``` Each parameter reaches the command as an environment variable: `BATCH_SIZE` arrives as `NODUS_PARAM_BATCH_SIZE`. Parameter names use uppercase letters, digits and `_`. ## Watch it and read the results [Section titled “Watch it and read the results”](#watch-it-and-read-the-results) ```console $ nodus get sweep/batch-size -w NAME PHASE CELLS COST BEST-COST AGE batch-size Succeeded 6/6 $0.31 2 14m $ nodus get sweep/batch-size -o yaml ``` Each cell runs as a Job named `-`, so `nodus logs job/batch-size-3` shows one cell’s output. The Sweep’s `status.cells` lists every cell with its GPU, region, parameters, phase, `costUSD`, `wallSeconds` and `unitsPerSecond`, and `status.best` names the cell with the lowest cost (`byCost`) and the highest throughput (`byThroughput`). `unitsPerSecond` counts the units your program reports with `nodus.log.unit(id, ms)` from the Python SDK, such as one call per batch or per request, divided by the cell’s wall time. Cells that report no units have no throughput. ## The matrix [Section titled “The matrix”](#the-matrix) | Field | Values | Limit | | ---------------- | ------------------------------------------------------------------------------------------------- | ------------------------ | | `matrix.gpu` | GPU requests, such as `L4`, `H100:8` or `H100!` | 16 | | `matrix.regions` | Region classes, such as `us` or `eu`; within the template’s `placement.regions` when it sets them | | | `matrix.params` | Parameter name → list of string values | 16 names, 64 values each | | `repetitions` | Runs of each combination, to measure variance | 64 | The number of cells is the product of the non-empty dimensions times `repetitions`, at most 256. Cells are numbered in a fixed order: GPU types vary slowest, then regions, then parameters by name, and repetitions fastest, so the repetitions of one combination sit next to each other. ## From Python [Section titled “From Python”](#from-python) ```python import nodus sweep = nodus.Sweep( {"image": "nodus/pytorch", "command": ["python", "bench.py"]}, # the Job spec every cell runs grid={"gpu": ["L4", "H100"], "BATCH_SIZE": [8, 16, 32]}, # gpu and region are dimensions; the rest are parameters repetitions=2, max_parallel=6, max_cost=25, ).run() report = sweep.wait() # a Sweep with failed cells is Failed (CellsFailed) and still has its report print(report.phase, report.best.by_cost) for cell in sweep.cells(): print(cell.index, cell.phase, cell.cost_usd, cell.wall_seconds) ``` `nodus.Sweep.from_name("batch-size")` reads one that exists, and `suspend()`, `resume()` and `cancel()` apply to every unfinished cell. A Function or a recipe as the target is not available yet and raises `nodus.errors.Unsupported`. ## Running and failing cells [Section titled “Running and failing cells”](#running-and-failing-cells) * `maxParallel` limits how many cells run at once (default 4). Cells start in index order. * `maxCostUSD` caps the whole Sweep. At the cap the running cells save their state and suspend, no new cell starts, and the Sweep becomes `Suspended` with reason `MaxCostReached`; raise the cap to continue. * A failed cell does not stop the others. When every cell has finished and some failed, the Sweep fails with reason `CellsFailed`, and its `status.cells` still carries the results of the cells that succeeded. * `nodus suspend sweep/x`, `resume` and `cancel` apply to every unfinished cell. Deleting a Sweep deletes its cell Jobs. # Training (Beta) > Fine-tune, post-train, distill and pretrain models with TrainingJobs on managed runtimes, on one GPU or across several machines. Beta TrainingJob and TrainingRuntime are `nodus.dev/v1beta1`. Fields may still change before they reach v1; a change is announced in the changelog with a migration note. Environments, used for RL and evaluation, are GA. A TrainingJob says what to train (a base model, a dataset or an environment, and the runtime’s parameters) and Nodus runs it as a Job: it caches the model, picks GPUs from the runtime’s presets, estimates the run before it starts, checkpoints it, and collects the trained weights as outputs. ## Sign in [Section titled “Sign in”](#sign-in) ```console $ pip install nodus-compute $ nodus login ``` ## Choose a runtime [Section titled “Choose a runtime”](#choose-a-runtime) A TrainingRuntime is a pinned trainer image with a parameter schema, presets and measured step times. The managed runtimes live in the `nodus` catalog project: | Runtime | What it trains | Data | | -------------------- | ------------------------------------------------------------------------------------------------------ | ------------------------------- | | `nodus/sft` | Supervised fine-tuning, full or LoRA/QLoRA; `task: Pretrain` for continued or from-scratch pretraining | prompt/completion or text JSONL | | `nodus/dpo` | Direct preference optimization | prompt, chosen, rejected | | `nodus/orpo` | Odds-ratio preference optimization (no reference model) | prompt, chosen, rejected | | `nodus/kto` | Kahneman-Tversky optimization from thumbs-up/down labels | prompt, completion, label | | `nodus/reward-model` | A reward model for RLHF | prompt, chosen, rejected | | `nodus/grpo-lora` | GRPO reinforcement learning against an Environment (single node) | an Environment | | `nodus/distill` | Knowledge distillation from a teacher model | prompts or chats | | `nodus/evaluate` | Evaluation only (`mode: Evaluate`) | an Environment or a benchmark | ```console $ nodus get trainingruntimes -n nodus $ nodus get trainingruntime/sft -n nodus -o yaml # the parameter schema, presets and step times ``` In the console, **Environments** lists the same runtimes under **Training templates**. You start a run from the SDK or the CLI; the console shows it under **Runs › Post-training** with its stages, curves, results, cost and outputs. ## Fine-tune with LoRA [Section titled “Fine-tune with LoRA”](#fine-tune-with-lora) sft.yaml ```yaml apiVersion: nodus.dev/v1beta1 kind: TrainingJob metadata: name: support-sft spec: runtime: nodus/sft model: uri: hf://Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca secret: hf-token # only for gated repositories data: volume: support-chats # a Volume holding chats.jsonl path: chats.jsonl format: JSONL columns: {prompt: question, completion: answer} parameters: steps: 2000 lora: {r: 16, alpha: 32} maxCostUSD: "20.00" ``` Review it before anything runs. A dry run compiles the TrainingJob into the Job it will create and prices it: ```console $ nodus apply -f sft.yaml --dry-run -o yaml # status.estimate: cost to completion, first hold, start time $ nodus apply -f sft.yaml ``` What the dry run decides: * **Resources** come from the first runtime preset that matches the model, the method (Full, LoRA or QLoRA, read from `parameters`) and quantization. Anything you set under `spec.resources` wins. * **The estimate** is the number of steps times the measured seconds per step on the slowest GPU type you allowed, plus evaluation tasks times seconds per task. Set `spec.expectedDuration` to use your own figure instead. * **Parameters** are merged over the runtime’s defaults and checked against its schema. A typo such as `learning_rate` is rejected with the field path, before you are charged anything. Model, teacher, and dataset Volume inputs are fixed to their latest committed revision when you create the TrainingJob, unless you select an explicit revision. Upload the files and wait for the Volume to be Ready first. Later uploads do not change the admitted run. Pin Hugging Face model revisions to a 40-character commit. The first TrainingJob that names a revision imports it once into a read-only Volume in your project, and every later run on that revision reuses it. The trainer runs with `HF_HUB_OFFLINE=1`, so billed GPU time is never spent downloading weights. ## Watch it [Section titled “Watch it”](#watch-it) ```console $ nodus get tj -w NAME RUNTIME PHASE STAGE CHANGE COST AGE support-sft nodus/sft Running Training $0.12 6m $ nodus describe tj/support-sft $ nodus logs job/support-sft -f ``` `STAGE` is `CachingModel` while the model imports, then the stage the trainer reports (`Baseline`, `Training`, `Evaluating`), and `Exporting` while the outputs are written to `spec.exportTo`. `describe` adds the step, total steps, loss and tokens per second, the provenance of the run (runtime and environment digests, model revision, and whether the runtime is a managed one) and the cost so far. In the console, open the run from **Compute**: the Overview shows the stages, step, loss and the baseline and final pass rates; **Gradings** lists every graded task as it is scored; **Results** has the outputs and the before and after table. The TrainingJob creates a Job of the same name and owns it. Logs, checkpoints and outputs are the Job’s, so everything in the [Jobs guide](/docs/guides/jobs/) applies. ## Download the results [Section titled “Download the results”](#download-the-results) Set `spec.exportTo.volume` to a Ready `ReadWriteOnce` Volume in the same project to also copy the complete result there. The runtime must declare an output at `/nodus/outputs`. The export uses an isolated CPU Job under the TrainingJob’s spending cap, waits up to 5 minutes for capacity, and allows 15 minutes for the copy. It follows suspension and cancellation. Success records the exact committed Volume revision in the `Exported` condition; files remain downloadable as Job outputs if the copy fails. Source or destination symbolic links are rejected. When the run succeeds, every file the trainer wrote under `/nodus/outputs` is an output of the TrainingJob, named by its path: `results.json` (the final metrics), `provenance.json`, `manifest.json` (a sha256 per file) and the weights under `adapter/` (LoRA and QLoRA) or `model/` (full fine-tuning), with the tokenizer beside them. ```console $ nodus get tj/support-sft -o jsonpath='{.status.outputs[*].name}' $ nodus cp tj/support-sft:outputs/results.json ./results.json $ nodus cp tj/support-sft:outputs/adapter/adapter_model.safetensors ./adapter_model.safetensors ``` `nodus cp` verifies each file against the sha256 recorded when it was collected. ## Full fine-tuning and pretraining [Section titled “Full fine-tuning and pretraining”](#full-fine-tuning-and-pretraining) Leave out the `lora` block for full fine-tuning; presets for the Full method pick larger GPUs. Pretraining uses the same runtime with `task: Pretrain`. `initialization: Continued` (the default) continues from the model’s weights; `Scratch` uses only its architecture and tokenizer: ```yaml spec: runtime: nodus/sft task: Pretrain model: {uri: hf://Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca} data: {volume: corpus, format: JSONL} parameters: {steps: 10000, initialization: Scratch, packing: true} resources: {gpu: {type: [h100-sxm-80g], count: 8}} ``` ## Preference tuning [Section titled “Preference tuning”](#preference-tuning) `dpo`, `orpo`, `kto` and `reward-model` take a dataset whose columns you map onto the names the runtime expects: ```yaml spec: runtime: nodus/dpo model: {volume: {name: my-sft-model, revision: 3}} # a model you trained earlier, by Volume data: volume: prefs format: JSONL columns: {prompt: prompt, chosen: chosen, rejected: rejected} parameters: {steps: 400} ``` A model can come from a Hugging Face URI or from one of your Volumes at a pinned revision, so a DPO run can start from the output of an SFT run. ## Distillation [Section titled “Distillation”](#distillation) `distill` trains the student in `model` on the teacher’s distributions. The teacher is cached the same way: ```yaml spec: runtime: nodus/distill model: {uri: hf://Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca} teacher: {uri: hf://Qwen/Qwen3-4B@0e9e39f249a16976918f6564b8830bc894c89659} data: {volume: chats} parameters: {temperature: 2.0} resources: {gpu: {type: [h100-sxm-80g], count: 2}} ``` ## Reinforcement learning (Beta, single node) [Section titled “Reinforcement learning (Beta, single node)”](#reinforcement-learning-beta-single-node) `grpo-lora` trains against an [Environment](/docs/guides/environments/): it samples completions for the environment’s tasks and Nodus grades them. The trainer never sees the expected answers and never grades itself: it submits completions to the TrainingJob’s `gradings` endpoint, and Nodus runs the environment’s grader in isolated containers on the same GPU host, with no network access, model files or GPU devices and the answers readable only by the grader. Task generation also uses this host; Nodus does not provision a separate CPU machine. A grader that fails is reported as an infrastructure failure and never scores the completion 0. Environment-backed training and evaluation require GPU capacity that supports isolated grading. Grading is serialized per TrainingJob. Baseline and final model evaluation run on the GPU. Custom runtime images must implement the task-fetch endpoint and declare `spec.environmentTasks: AttemptAPI`; older runtimes are rejected before GPU acquisition. ```yaml spec: runtime: nodus/grpo-lora model: {uri: hf://Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca} environment: {name: nodus/graph-coloring@1.0.0, trainTasks: 50, heldOutTasks: 64, seed: 42} evaluation: {baseline: true, final: true} parameters: {steps: 50, lora: {r: 8}} resources: {gpu: {type: [rtx4090-24g], count: 1}} maxCostUSD: "5.00" ``` With `evaluation.baseline` and `evaluation.final`, the run scores the held-out tasks before and after training. `status.summary` holds both pass rates and a comparison: the change in percentage points, and whether it is comparable at all. A comparison needs the same task set on both sides, at least 16 scored tasks, and a result that is not all-pass or all-fail on both sides; otherwise its `reason` says why (`TooFewTasks`, `DifferentTasks`, `NonePassed`, `AllPassed`, `IncompleteEvents`). `status.diagnostics.rewardSignal` warns when every training reward was the same, which teaches a group-relative method nothing. ### Reproducing the legacy letter-counting protocol [Section titled “Reproducing the legacy letter-counting protocol”](#reproducing-the-legacy-letter-counting-protocol) Use `nodus/letter-counting-legacy-eval@1.0.0` for the legacy evaluation protocol: the first 64 tasks from Reasoning Gym’s `letter_counting` generator with seed 42, an answer-only system message, greedy decoding with 256 completion tokens, and the generator’s raw-completion scorer. The `letter-counting` catalog example pins Qwen3-1.7B and generates one prompt at a time. Use `nodus/letter-counting-legacy-rl@1.0.0` for the legacy RL protocol. Its first 64 tasks are training data; the following 64 are held out for both baseline and final evaluation. The example sets 50 optimizer steps, learning rate `5e-6`, eight generations, generation batch size eight, per-GPU batch size one, gradient accumulation four, temperature `1.2`, top-p `0.95`, top-k `50`, and 64 completion tokens. It uses a rank-eight LoRA adapter with alpha 16, no dropout, all linear layers, bf16, non-reentrant gradient checkpointing, and a linear learning-rate schedule without warmup. Thinking is disabled in both phases. Recovery checkpoints remain enabled. These profiles pin Reasoning Gym `0.1.25` and preserve the legacy source’s task slicing, prompt roles and scoring. They do not establish historical binary equivalence or a measured pass rate. The evaluation profile’s 64 tasks are different from the RL profile’s held-out 64 tasks; compare baseline and final within the same RL run. The existing `nodus/reasoning-gym` profile uses shuffled family-specific seeds and remains a different benchmark. Record the exact runtime, environment and model digests alongside results before claiming a numeric comparison. RL runs on one node. Multi-node RL and custom environment collections are not available yet. ## Multi-node training [Section titled “Multi-node training”](#multi-node-training) Set `resources.nodes` above 1 to train across several machines, each with `resources.gpu.count` GPUs. The TrainingJob compiles into one multi-node Job (a gang): every node starts together, `torchrun` is launched on each with its rank and the rendezvous address, and the run checkpoints with `torch.distributed.checkpoint`. ```yaml spec: runtime: nodus/dpo model: {volume: {name: my-8b, revision: 3}} data: {volume: prefs, format: JSONL, columns: {prompt: prompt, chosen: chosen, rejected: rejected}} parameters: {steps: 400} resources: {gpu: {type: [h100-sxm-80g], count: 8}, nodes: 2} distributed: {network: Global, transport: Auto} ``` * **Launchers.** A runtime declares `Torchrun`, `Accelerate`, `Deepspeed` or `Plain`, and the run uses that launcher on every node: `torchrun` reads the gang’s rendezvous settings, `accelerate launch` and DeepSpeed’s per-node launcher read the rank, address and world size Nodus sets on each node, and `Plain` starts the runtime’s command once per node. A single node runs `torchrun --standalone` for a `Torchrun` runtime, `accelerate launch` or `deepspeed` for the other two. * **Placement.** `distributed.network` is `Colocated` (one provider and region, the fastest interconnect), `Regional` or `Global`; `transport: Auto` also allows an encrypted SSH connection between two restricted GPU machines without a separate relay server. Native GPU-to-GPU connectivity is preferred. Wider settings find capacity sooner at some cost in throughput. * **Recovery.** If a node is lost, the whole gang restarts from the latest checkpoint every rank committed, up to 3 times; progress since that checkpoint is redone. A single-node run restores its latest snapshot instead. Managed trainers bind recovery state to the full run configuration, runtime image digest, and pinned model and dataset references; incompatible state is refused. GRPO reuses the baseline saved in that state so recovery does not repeat baseline grading. * **Cost.** The estimate covers every node. Assembly time (nodes waiting for each other) is billed and bounded; see [Multi-node training](/docs/guides/multi-node/) for the limits. ## Spending caps [Section titled “Spending caps”](#spending-caps) `spec.maxCostUSD` caps the whole TrainingJob: its Job, its same-host grading, its output export and its retries. When the cap is reached the run is suspended, not failed, and active grader containers are removed. Raise the cap to resume the same run from its checkpoint: ```console $ nodus patch tj/support-sft --type merge -p '{"spec":{"maxCostUSD":"40"}}' ``` The cap can only be raised. `spec.state: Suspended` pauses a run yourself, and `Running` resumes it. ## Tracking [Section titled “Tracking”](#tracking) `spec.tracking.mlflow` (`uri`, `secret`, `experiment`) or `spec.tracking.wandb` (`connection`) send the trainer’s metrics to your own MLflow server or Weights & Biases project as well. Nodus shows its own stage and metrics either way. The catalog runtimes report to MLflow when the run has `MLFLOW_TRACKING_URI` and to W\&B when it has `WANDB_API_KEY`, which is what those two fields set; the Secret holds `MLFLOW_TRACKING_TOKEN` (or `MLFLOW_TRACKING_USERNAME` and `MLFLOW_TRACKING_PASSWORD`), and the W\&B Connection’s Secret holds `WANDB_API_KEY`. ## Bring your own runtime [Section titled “Bring your own runtime”](#bring-your-own-runtime) A TrainingRuntime in your project can run any image. The trainer reads: | Variable | Contents | | -------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `NODUS_INPUT_PARAMS` | a directory that holds `params.json`, the run: `task` and `mode` (for example `SFT`, `Train`), the merged, schema-checked `parameters`, `data` (path, format and column mapping), `environment` (task counts, seed and graders), `evaluation`, `trainingJob` and the pinned `modelRevision` and `teacherRevision` | | Environment tasks | Call `POST …/trainingjobs/{name}/tasks` through the attempt API socket. The response `files` maps `train.jsonl` and `test.jsonl` to JSONL strings containing `{taskId, prompt, metadata}` and no hidden answers. Managed runtimes fetch them automatically; custom runtimes must implement this step. | | `NODUS_INPUT_MODEL`, `NODUS_INPUT_TEACHER`, `NODUS_INPUT_DATA` | where the model, teacher and data are mounted | | `NODUS_STATE_DIR`, `NODUS_OUTPUT_DIR` | resumable state and results | | `NODUS_NODE_RANK`, `MASTER_ADDR`, `PET_*` | the gang of a multi-node run | | `NODUS_CHECKPOINT_URI`, `NODUS_RESTORE_URI` | where ranks write and restore distributed checkpoints | Write and load your own model, optimizer and progress files under the state directory; Nodus decides when to checkpoint and where checkpoints are stored, and never adds resume flags to your command. # Volumes > Keep datasets, model weights and working files in Volumes, upload and download them, import from Hugging Face, git, URLs or S3, and mount them in Jobs and Sandboxes. A Volume is named storage that outlives the work that uses it. Every change is saved as a numbered revision, so you can see what changed, mount an earlier revision read-only, and never lose files to a stopped machine. You pay for the bytes stored, after your org’s included 10 GB. ## Create a Volume and upload files [Section titled “Create a Volume and upload files”](#create-a-volume-and-upload-files) volume.yaml ```yaml apiVersion: nodus.dev/v1 kind: Volume metadata: name: ex-volumes-put-get spec: accessMode: ReadWriteOnce size: 1Gi source: {upload: {}} ``` ```console nodus apply -f volume.yaml # or: nodus create volume data --size 100Gi nodus volume put ex-volumes-put-get ./data /data # uploads ./data and prints the new revision nodus volume ls ex-volumes-put-get /data nodus volume get ex-volumes-put-get /data/hello.txt ./out/ nodus volume put data ./archive.tar.gz /raw --extract nodus volume rm data /raw/old.csv nodus volume clear data # an empty revision; the Volume and its history stay ``` Uploads only send what changed: a second `put` of the same files moves almost nothing. An upload that started from an older revision than the latest is refused with `409 Conflict`, so two people uploading at once never overwrite each other silently; run it again on top of the new revision. `put` adds to what is there: a directory’s contents merge into the remote path, and a file lands at the remote path, or inside it when the remote path is a directory or ends in `/`. `rm` and `clear` also commit new revisions, so earlier revisions keep the removed files. `get` and `ls` read the latest revision, or the one `--revision ` names; `get … -` prints a file to standard output, and existing local files are kept unless `--force`. `--extract` expands a `.zip`, `.tar`, `.tar.gz`, `.tgz` or `.tar.zst` archive. It refuses an archive whose paths or links lead outside the remote path, or that has more than 10,000 members. ## Mount it in a Job [Section titled “Mount it in a Job”](#mount-it-in-a-job) job.yaml ```yaml apiVersion: nodus.dev/v1 kind: Job metadata: name: ex-volumes-put-get spec: image: nodus/python:3.12 volumes: - {volume: ex-volumes-put-get, mountPath: /mnt/vol, readOnly: true} command: [cat, /mnt/vol/data/hello.txt] ``` | `accessMode` | Who can write | How changes are saved | | ------------------------------------ | ----------------------------- | -------------------------------------------------------------------------------------------- | | `ReadWriteOnce` (default) | One attempt at a time | On exit and every `commitInterval` (default 5 minutes) | | `ReadOnlyMany` (default for imports) | Nobody; any number of readers | Revisions come from imports and uploads | | `ReadWriteMany` | Every worker | Each worker’s changes are published when it commits or exits; the last writer of a path wins | A `ReadWriteOnce` Volume is held by one attempt at a time. Starting a second writer, or deleting the Volume while it is held, fails with `409 VolumeBusy`, which names the holder; mount it with `readOnly: true` to read alongside. Functions and Agents with more than one worker need `ReadOnlyMany` or `ReadWriteMany`. Mount an earlier revision with `revision: ` and `readOnly: true`. `nodus get volume data` shows the latest `REVISION`, `USED` and the current `HOLDER`; `GET …/volumes/{name}/revisions` lists the kept revisions (`revisionHistoryLimit`, default 10). ## Import from elsewhere [Section titled “Import from elsewhere”](#import-from-elsewhere) Set exactly one `source`; the import runs once on Nodus, billed as CPU time, and produces revision 1: ```yaml spec: accessMode: ReadOnlyMany size: 200Gi maxCostUSD: "2.00" source: {huggingface: {repo: meta-llama/Llama-3.1-8B, revision: 0e9e39f249a16976918f6564b8830bc894c89659, secret: hf-token}} ``` | Source | Fields | | ------------- | ------------------------------------------------------------------------------------------------------- | | `huggingface` | `repo`, `revision` (a commit or tag), `files` globs, `secret` holding `HF_TOKEN` | | `git` | `repo`, `ref`, `lfs` | | `url` | an `https://` `url`, `sha256`, `extract: Auto` unpacks `.zip`, `.tar`, `.tar.gz`, `.tgz` and `.tar.zst` | | `s3` | `uri` and an `S3` [Connection](/docs/guides/connections/) | The Volume is `Pending` while it imports and `Ready` after. A failed import shows `Failed` with reason `ImportFailed` and a message; fix the source or add credits, then run `nodus request reimport volume/`. ## In Python [Section titled “In Python”](#in-python) ```python vol = nodus.Volume.from_name("data", create_if_missing=True) vol.put_file("./local.csv", "/train/local.csv") print(vol.listdir("/train")) ``` # Webhooks > Get signed HTTPS callbacks when Jobs finish, balances run low and other events happen, and verify them. A **WebhookEndpoint** sends an HTTPS request to your server each time an event you care about happens: a Job succeeds or fails, a balance runs low, a Budget crosses a threshold. You can have as many endpoints as you need, each with its own URL, event types and projects. Every delivery is signed with the [Standard Webhooks](https://www.standardwebhooks.com) scheme, so you can check that it came from Nodus and was not changed in transit. ## Create an endpoint [Section titled “Create an endpoint”](#create-an-endpoint) endpoint.yaml ```yaml apiVersion: nodus.dev/v1 kind: WebhookEndpoint metadata: name: example-receiver spec: url: https://hooks.example.com/nodus description: Job results for the example receiver eventTypes: - job.succeeded - job.failed - billing.low_balance ``` ```console $ nodus apply -f endpoint.yaml -o json | jq -r .status.signingSecret whsec_... ``` The signing secret is in the response to the create, and nowhere else. Copy it into your receiver’s configuration now. If you lose it, [rotate it](#rotate-the-signing-secret) to get a new one. Nodus sends a `webhookendpoint.ping` delivery right after the create, so you can see your receiver work before any real event. | Field | Meaning | | ------------- | -------------------------------------------------------------------------------------------------------------- | | `url` | Where deliveries go. `https` only, with a public address. No redirects are followed. | | `eventTypes` | Event types to send: an exact type (`job.succeeded`), a prefix (`job.*`, `billing.*`) or `*`. Defaults to `*`. | | `projects` | Send project events only for these projects. Empty means every project. Org-level events always match. | | `enabled` | `false` pauses deliveries. Pending deliveries to a paused endpoint are dropped. | | `description` | A note for you. | Every field can change after the create: ```console $ nodus patch webhookendpoint ci --type merge -p '{"spec":{"eventTypes":["job.*"]}}' ``` ## What Nodus sends [Section titled “What Nodus sends”](#what-nodus-sends) A delivery is a `POST` with a JSON body: ```json { "type": "job.succeeded", "timestamp": "2026-10-01T12:00:00Z", "data": { "event": { "reason": "Succeeded", "message": "…", "count": 1 }, "object": { "apiVersion": "nodus.dev/v1", "kind": "Job", "metadata": { "name": "train", "namespace": "default", "uid": "job_…" }, "status": { "phase": "Succeeded" } } } } ``` and three headers: | Header | Meaning | | ------------------- | -------------------------------------------------------------------------------------- | | `webhook-id` | Names this delivery. It stays the same across retries, so use it to ignore duplicates. | | `webhook-timestamp` | When the attempt was signed, in Unix seconds. | | `webhook-signature` | One or more signatures, separated by spaces, each like `v1,`. | The event type is `.` in lowercase with underscores, such as `job.succeeded` or `agentrun.needs_resolution`. Billing events use the `billing.` prefix. ## Event types [Section titled “Event types”](#event-types) | Type | When | | ----------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------- | | `job.succeeded`, `job.failed`, `job.cancelled` | A Job reached a final phase. The same reasons exist for other run kinds, such as `agentrun.failed`. | | `job.preempted`, `job.node_lost`, `job.restored` | A Job lost its machine, and was restored from its last checkpoint. | | `job.checkpointed`, `job.checkpoint_failed` | A checkpoint was saved, or could not be. | | `job.max_cost_reached`, `job.funding_lost`, `job.funding_restored` | A Job stopped or resumed because of its cost cap or your balance. The same reasons exist for Sandboxes and other kinds with a cost. | | `job.output_sink_failed` | A Job’s output could not be loaded into its destination. | | `job.gang_restarted`, `job.network_degraded` | A multi-node Job restarted or lost network quality. | | `sandbox.started`, `sandbox.stopped`, `sandbox.stuck`, `sandbox.setup_failed` | A Sandbox changed state. | | `agentrun.needs_resolution` | An AgentRun is waiting for you to resolve a step. | | `billing.low_balance` | Your available balance fell below the warning level. | | `billing.auto_recharge_succeeded`, `billing.auto_recharge_failed` | An auto-recharge added credits, or its charge failed. | | `billing.arrears_posted` | A charge was larger than your balance, so new work waits until it is paid. | | `billing.dispute_opened`, `billing.dispute_closed` | A payment was disputed, which pauses new work, or the dispute closed. | | `billing.storage_quota_exceeded`, `billing.egress_quota_exceeded` | A quota was reached. | | `budget.threshold`, `budget.exceeded` | A Budget crossed a threshold, or refused a new run. | | `topup.succeeded`, `topup.failed`, `topup.refunded` | A top-up completed, its payment failed, or it was refunded. | | `pool.node_offline`, `pool.forecast_shortfall`, `pool.action_proposed` | Events of your own capacity pools. | | `webhookendpoint.ping` | Sent once when an endpoint is created. | | `webhookendpoint.deliveries_failing` | Your receiver has failed every recent delivery. | `nodus get events` shows the events of your org; each webhook type is the kind and reason of one of them. ## Verify the signature [Section titled “Verify the signature”](#verify-the-signature) Check every delivery before you trust it. The signature covers the exact bytes of the body, so verify the raw body before you parse it. Use a Standard Webhooks library for your language; it also rejects timestamps more than five minutes old, which stops a captured delivery from being replayed later. main.go ```go package main import ( "encoding/json" "fmt" "io" "log" "net/http" "os" "sync" "time" standardwebhooks "github.com/standard-webhooks/standard-webhooks/libraries/go" ) // maxBody bounds a delivery; Nodus sends at most 64 KiB. const maxBody = 64 << 10 // event is the body of a delivery: the event's type and the Event and object it is about. type event struct { Type string `json:"type"` Timestamp time.Time `json:"timestamp"` Data struct { Event json.RawMessage `json:"event"` Object struct { Kind string `json:"kind"` Metadata struct { Name string `json:"name"` Namespace string `json:"namespace"` } `json:"metadata"` } `json:"object"` } `json:"data"` } // receiver verifies and handles deliveries. type receiver struct { hook *standardwebhooks.Webhook out io.Writer mu sync.Mutex seen map[string]bool } func newReceiver(secret string, out io.Writer) (*receiver, error) { hook, err := standardwebhooks.NewWebhook(secret) if err != nil { return nil, fmt.Errorf("NODUS_WEBHOOK_SECRET: %w", err) } return &receiver{hook: hook, out: out, seen: map[string]bool{}}, nil } // ServeHTTP answers 2xx once the delivery is verified and recorded, which is all Nodus needs to stop retrying. A // delivery that fails verification gets 400 and is never parsed: the signature covers the exact bytes received, so // verify before you decode. func (r *receiver) ServeHTTP(w http.ResponseWriter, req *http.Request) { body, err := io.ReadAll(http.MaxBytesReader(w, req.Body, maxBody)) if err != nil { http.Error(w, "body too large", http.StatusRequestEntityTooLarge) return } // Verify checks the signature against every secret in the header and rejects timestamps more than five minutes // from now, so a captured delivery cannot be replayed later. if err := r.hook.Verify(body, req.Header); err != nil { http.Error(w, "invalid signature", http.StatusBadRequest) return } // A delivery can arrive more than once (a retry after a slow answer): webhook-id names the delivery, so handle // each id once and still answer 200. if r.firstTime(req.Header.Get(standardwebhooks.HeaderWebhookID)) { var e event if err := json.Unmarshal(body, &e); err != nil { http.Error(w, "malformed event", http.StatusBadRequest) return } r.handle(e) } w.WriteHeader(http.StatusOK) } func (r *receiver) firstTime(id string) bool { r.mu.Lock() defer r.mu.Unlock() if r.seen[id] { return false } r.seen[id] = true return true } // handle does the work of an event. Do slow work after answering (queue it); Nodus waits 15 seconds for a response. func (r *receiver) handle(e event) { obj := e.Data.Object fmt.Fprintf(r.out, "%s %s/%s\n", e.Type, obj.Kind, obj.Metadata.Name) } func main() { rcv, err := newReceiver(os.Getenv("NODUS_WEBHOOK_SECRET"), os.Stdout) if err != nil { log.Fatal(err) } srv := &http.Server{Addr: ":8080", Handler: rcv, ReadHeaderTimeout: 5 * time.Second} log.Fatal(srv.ListenAndServe()) } ``` This receiver runs as is: `NODUS_WEBHOOK_SECRET=whsec_... go run ./examples/webhooks/receiver`. The same check in Python: ```python from standardwebhooks.webhooks import Webhook wh = Webhook(secret) # the whsec_ secret event = wh.verify(request_body, request_headers) # raises if the signature or timestamp is wrong ``` If you cannot use a library, compute `base64(HMAC-SHA256(key, id + "." + timestamp + "." + body))`, where `key` is the secret after the `whsec_` prefix, base64-decoded, and compare it in constant time with each `v1,` signature in the header. Caution Answer with a `2xx` status as soon as you have verified and recorded the delivery, and do the slow work after. Nodus waits 15 seconds for a response, and treats a timeout like a failure. ## Retries and delivery order [Section titled “Retries and delivery order”](#retries-and-delivery-order) A `2xx` response is success. Anything else, a timeout or a connection error is retried with exponential backoff, starting at 30 seconds and doubling up to 8 hours between attempts, for 3 days after the event. After that the delivery is marked failed. A URL that Nodus is not allowed to connect to, such as a private address, fails at once. Deliveries can arrive more than once and out of order. Use `webhook-id` to ignore a duplicate, and the `timestamp` in the body or the object’s `status` to see which state is newer. ## The delivery log [Section titled “The delivery log”](#the-delivery-log) Nodus keeps the log of an endpoint’s deliveries for 7 days: ```console $ curl -s "$NODUS_API_URL/apis/nodus.dev/v1/webhookendpoints/ci/deliveries?limit=20" \ -H "Authorization: Bearer $NODUS_API_KEY" ``` Each entry shows the event type, the number of attempts, the receiver’s last response code, how long it took, and a `phase`: `Pending` while retries remain, then `Succeeded` or `Failed`. A delivery to an endpoint that was deleted or disabled is `Abandoned`. The response includes `continue` when there is more; pass it back as `after`. To send a delivery again, replay it. The replay has the same body and a new `webhook-id`: ```console $ curl -s -X POST "$NODUS_API_URL/apis/nodus.dev/v1/webhookendpoints/ci/deliveries/$ID/replays" \ -H "Authorization: Bearer $NODUS_API_KEY" ``` ## Rotate the signing secret [Section titled “Rotate the signing secret”](#rotate-the-signing-secret) Rotate when a secret may have leaked, or when you lost it: ```console $ curl -s -X POST "$NODUS_API_URL/apis/nodus.dev/v1/webhookendpoints/ci/rotations" \ -H "Authorization: Bearer $NODUS_API_KEY" -H "Content-Type: application/json" \ -H "Idempotency-Key: $(uuidgen)" -d '{"previousSecretTTL": "24h"}' | jq -r .status.signingSecret ``` The response holds the new secret, once. For `previousSecretTTL` (24 hours by default, at most 7 days) every delivery carries a signature for the new secret and one for the old, so a receiver using either verifies it. Switch your receiver to the new secret, and the old one stops signing when the time is up. `status.secretRotation` shows when. ## Check an endpoint’s health [Section titled “Check an endpoint’s health”](#check-an-endpoints-health) ```console $ nodus get webhookendpoint ci -o yaml ``` `status` shows the latest delivery (`lastDelivery`), the deliveries of the last 24 hours that succeeded or failed (`deliveries24h`), and, when the last three deliveries in a row failed, `failingSince` with the `Healthy` condition set to `False` and reason `DeliveriesFailing`. Nodus also records a `webhookendpoint.deliveries_failing` event when that happens, which you can send to another endpoint. Note Webhooks never carry provider names, machine ids or other internal details. Their bodies hold what the API would show you for the same object. # Workspaces > A GPU or CPU development machine with a saved home directory, reached with SSH, VS Code or JupyterLab. A Workspace is your own development machine: a GPU or CPU machine with VS Code in the browser, JupyterLab and SSH. Its home directory, `/home/nodus`, lives on a Volume and is saved every time the Workspace stops, so you can stop it at the end of the day and pick up where you left off. While it is stopped you pay only for the saved files. You need the `nodus` CLI to connect over SSH ([install it](/docs/getting-started/) and run `nodus login`). Browser VS Code and JupyterLab need only the console. ## Create a Workspace [Section titled “Create a Workspace”](#create-a-workspace) ```console $ nodus create workspace lab --gpu H100 $ nodus get workspaces NAME PHASE COMPUTE HOME STOP-REASON COST AGE lab Running 1 × H100 lab-home $0.09 3m ``` `--gpu H100:2` asks for two GPUs; `--cpu 8 --memory 32Gi` without `--gpu` creates a CPU Workspace. The home Volume is `lab-home`, created on the first start if it does not exist; set `spec.volume` to use another one. The default image is `nodus/workspace-pytorch-cuda` (PyTorch and CUDA) on NVIDIA GPUs, `nodus/workspace-rocm` on AMD GPUs and `nodus/workspace-cpu` on CPU. Every catalog image has VS Code, JupyterLab, Python and the user `nodus`. The same Workspace as a manifest, applied with `nodus apply -f workspace.yaml`: workspace.yaml ```yaml apiVersion: nodus.dev/v1 kind: Workspace metadata: name: ws-home spec: image: nodus/workspace-cpu resources: cpu: "2" memory: 8Gi volume: ws-home-home idleTimeout: 30m ``` In the console, **Workspaces › New workspace** shows the GPU picker, the estimate and the same manifest as YAML, CLI and Python before you create it. ## Connect [Section titled “Connect”](#connect) ### SSH [Section titled “SSH”](#ssh) ```console $ nodus login $ nodus ssh workspace/lab nodus@lab:~$ nvidia-smi $ nodus ssh workspace/lab -- python train.py --epochs 1 ``` Sign in once with `nodus login` on each computer using an up-to-date CLI and OpenSSH 8.5 or newer. Login creates a dedicated local SSH key for your account, registers only its public key, and configures every existing and future CPU or GPU Workspace in the organizations you selected. New Workspaces need no separate setup. Your private key stays on your computer with owner-only permissions; sign-in does not start or connect to a Workspace. Each teammate signs in to their own account. Organization roles and project access determine which Workspaces they can enter. Remove a device’s key under **Account › SSH keys** to revoke it. Removing a teammate’s access also prevents their keys from opening that organization’s Workspaces. `nodus ssh` connects through the Nodus gateway. No port is open on the machine. OpenSSH verifies the Workspace’s host key against the authenticated API on every connection, with strict host-key checking. A missing or revoked identity fails closed; it does not trust a host key on first use. ### VS Code Remote-SSH, Cursor and JetBrains Gateway [Section titled “VS Code Remote-SSH, Cursor and JetBrains Gateway”](#vs-code-remote-ssh-cursor-and-jetbrains-gateway) After `nodus login`, open **Desktop VS Code** or **Cursor** in the console, or select the Workspace host in your editor’s Remote-SSH connection dialog. The same account key works for all your authorized Workspaces, including ones created later. Signing in to the website alone cannot configure SSH on your computer. The `.nodus` host is an SSH alias, not a public DNS address. Login records the CLI executable and saved context, so editors do not depend on your terminal’s environment. Plain `ssh`, `scp` and `rsync` use the same connection. `nodus ssh workspace/lab --config` prints the rule; `nodus ssh workspace/lab --setup` is an optional repair command. Neither command starts or connects to a Workspace. ### Browser VS Code and JupyterLab [Section titled “Browser VS Code and JupyterLab”](#browser-vs-code-and-jupyterlab) Open the Workspace in the console and choose **Browser VS Code** or **JupyterLab**. Both open on a private address that only members of your organization can reach after signing in. ## Stop, start and what is saved [Section titled “Stop, start and what is saved”](#stop-start-and-what-is-saved) ```console $ nodus stop workspace/lab $ nodus start workspace/lab ``` When a Workspace stops, everything in `/home/nodus` is saved to its home Volume and the machine is released. Everything else is temporary, including the `/workspace` directory (the container’s working directory, which has nothing to do with the Workspace kind) and anything installed outside your home. Install Python packages with `pip install --user` or in a virtual environment under `/home/nodus` to keep them. The next start restores the saved home on a fresh machine. Use `ephemeral: true` for a Workspace with no home Volume, whose files are deleted on stop. A Workspace also stops by itself: | Stop reason | When | How it starts again | | --------------------------------------------------------- | -------------------------------------------------------------------------- | ------------------------------------------------------------- | | `Idle` | No SSH session, editor traffic or process for `idleTimeout` (default `1h`) | Connect, open a tool, or `nodus start` | | `Schedule` | At `schedule.stopAt` | The next `schedule.readyBy`, connecting, or `nodus start` | | `InsufficientCredits`, `BudgetExceeded`, `MaxCostReached` | Money ran out or a limit was reached | Add credits or raise the limit, then connect or `nodus start` | | `User` | You ran `nodus stop` or chose **Stop** | Only `nodus start` or **Start** | Connecting to a Workspace that stopped by itself starts it; `nodus ssh` waits and connects when it is ready. The console asks before a connect starts a stopped Workspace and shows the rate it will bill at. ### Schedules [Section titled “Schedules”](#schedules) ```yaml spec: schedule: readyBy: "2026-10-01T09:00:00-07:00" stopAt: "2026-10-01T19:00:00-07:00" ``` The Workspace starts 15 minutes before `readyBy` so it is ready on time, and stops at `stopAt`. It records a `ScheduledStart` event, or `ScheduleMissed` when it could not start in time. ## Run a job on your Workspace files [Section titled “Run a job on your Workspace files”](#run-a-job-on-your-workspace-files) ```console $ nodus run --from workspace/lab -- python train.py ``` The Job runs on a copy of the last saved home, labelled `nodus.dev/workspace=lab`, so the Workspace’s **Jobs from this workspace** list shows it. ## Sessions and billing [Section titled “Sessions and billing”](#sessions-and-billing) Each period a Workspace runs is a session, recorded as an Attempt with its receipt: ```console $ nodus get attempts -l nodus.dev/workspace=lab ``` * A GPU Workspace bills at the rate shown when it starts, from when its machine is acquired until the machine is released. A session that stops for any reason closes its charge. * A CPU Workspace bills at the listed CPU rates while it runs. * Saved files in the home Volume bill as storage after your organization’s included storage. `maxCostUSD` stops and saves the Workspace when its cost reaches the limit; raising the limit lets it start again. ## Shared access [Section titled “Shared access”](#shared-access) A Workspace belongs to its project. Every member of the organization with access to the project can see it, start it and connect with their own account and SSH key; `spec.sshKeys` limits SSH to the keys you name. ## When a Workspace fails [Section titled “When a Workspace fails”](#when-a-workspace-fails) A Workspace in the `Failed` phase does not restart. Delete it and create it again with the same name (or the same `spec.volume`): the home Volume keeps your files. ```console $ nodus delete workspace/lab $ nodus create workspace lab --gpu H100 ``` To delete the saved files too, delete the home Volume (`nodus delete volume/lab-home`), or empty it with `nodus volume clear lab-home`. # Reference > Look up exact CLI commands, Python signatures, API fields, errors, and prices. Use reference pages to check a specific detail. For a walkthrough, [find a guide](/docs/guides/) first. | I need… | Reference | | -------------------------------------- | --------------------------------------------------------------------- | | Command syntax and flags | [CLI](/docs/reference/cli/) | | Python classes, methods, and arguments | [Python SDK](/docs/reference/python/) | | HTTP endpoints and request fields | [HTTP API](/docs/reference/api/) · [OpenAPI JSON](/docs/openapi.json) | | An error’s meaning and fix | [Error codes](/docs/reference/errors/) | | Help diagnosing a failure | [Troubleshooting](/docs/reference/troubleshooting/) | | Compute, storage, and transfer prices | [Pricing](/docs/reference/pricing/) | | Training runtime options | [Training runtimes (Beta)](/docs/reference/runtimes/) | | Training and evaluation environments | [Environments](/docs/reference/environments/) | | Server configuration variables | [Server configuration](/docs/reference/server-configuration/) | Command, SDK, API, error, catalog, and price details are generated from their product sources. # HTTP API > The Nodus REST API, Kubernetes-style resources under /apis/nodus.dev, generated from the Go types. Nodus serves every resource at a Kubernetes-style path, so `kubectl`, `client-go` and the generated clients all work against it: ```text GET /apis/nodus.dev/v1/namespaces/{project}/jobs POST /apis/nodus.dev/v1/namespaces/{project}/jobs GET /apis/nodus.dev/v1/namespaces/{project}/jobs/{name} GET /apis/nodus.dev/v1/namespaces/{project}/jobs?watch=true ``` Authenticate with an API key as a bearer token (`Authorization: Bearer nodus_sk_…`). Lists page with `limit` and `continue`, writes accept an `Idempotency-Key` header, and `?dryRun=All` validates a create and returns its estimate without charging anything. Errors are described in [Error codes](/docs/reference/errors/). The per-operation reference in the sidebar is generated from the OpenAPI document, which is also served at [`/docs/openapi.json`](/docs/openapi.json). # CLI reference > Every nodus command, flag and example, generated from the CLI's command tree. The `nodus` CLI follows kubectl’s verbs (`get`, `describe`, `apply`, `delete`, `logs`, `exec`, `cp`) over every Nodus resource, plus task commands such as `run`, `login` and `billing`. Global flags: | Flag | Meaning | | ----------------------------------------------------- | -------------------------------------------------------------------------------------------- | | `--context`, `--org` | Select a context or org | | `-p, --project` (`-n` accepted), `-A, --all-projects` | Project scope | | `-o, --output` | `table`, `wide`, `json`, `yaml`, `name`, `jsonpath=…`, `custom-columns=…`, `csv`, `estimate` | | `-l, --selector`, `--field-selector` | Label and field selectors | | `-w, --watch` | Keep watching after listing | | `--dry-run=server` | Validate and return the estimate without creating anything | `nodus run` exits with your command’s exit code. Nodus and API errors exit `125` and timeouts `124`. ## Commands [Section titled “Commands”](#commands) [nodus ](/docs/reference/cli/nodus/)Run compute Jobs on Nodus [nodus agentrun ](/docs/reference/cli/nodus_agentrun/)Message an AgentRun and read its steps and answer [nodus agentrun answer ](/docs/reference/cli/nodus_agentrun_answer/)Print an AgentRun's last answer in full [nodus agentrun send ](/docs/reference/cli/nodus_agentrun_send/)Send a message to an AgentRun, which takes it as its next prompt [nodus agentrun steps ](/docs/reference/cli/nodus_agentrun_steps/)List an AgentRun's journal steps: what it did, in order [nodus annotate ](/docs/reference/cli/nodus_annotate/)Set or remove annotations [nodus api-resources ](/docs/reference/cli/nodus_api-resources/)List the kinds the server offers [nodus apply ](/docs/reference/cli/nodus_apply/)Create or update resources from manifests (client-side three-way merge) [nodus attach ](/docs/reference/cli/nodus_attach/)Attach to a running Process of a Job, Sandbox or Workspace [nodus auth ](/docs/reference/cli/nodus_auth/)Credentials for other tools, and permission checks [nodus auth can-i ](/docs/reference/cli/nodus_auth_can-i/)Check whether the current credential may perform an action [nodus auth token ](/docs/reference/cli/nodus_auth_token/)Print an ExecCredential for kubectl and client-go (the kubeconfig exec plugin) [nodus billing ](/docs/reference/cli/nodus_billing/)Show the org's balance; top up, redeem a code and manage auto-recharge [nodus billing auto-recharge ](/docs/reference/cli/nodus_billing_auto-recharge/)Top up automatically when available credit falls below a threshold [nodus billing deal ](/docs/reference/cli/nodus_billing_deal/)View partner offers or request an eligibility review [nodus billing portal ](/docs/reference/cli/nodus_billing_portal/)Open the Stripe Customer Portal: cards, invoices and billing details [nodus billing receipts ](/docs/reference/cli/nodus_billing_receipts/)List top-ups with their Stripe receipt links [nodus billing redeem ](/docs/reference/cli/nodus_billing_redeem/)Redeem a promo code for credit [nodus billing top-up ](/docs/reference/cli/nodus_billing_top-up/)Buy prepaid credit with Stripe Checkout ($5–$1,000) [nodus billing usage ](/docs/reference/cli/nodus_billing_usage/)Show metered usage, itemized or summed by project, kind, member, label, meter, day, segment, offering or rank [nodus cancel ](/docs/reference/cli/nodus_cancel/)Cancel a run, FunctionCall or Process [nodus cloud ](/docs/reference/cli/nodus_cloud/)Inspect connected cloud accounts [nodus cloud inventory ](/docs/reference/cli/nodus_cloud_inventory/)List the instances, GPUs or spend a cloud account's last sync saw [nodus completion ](/docs/reference/cli/nodus_completion/)Generate the autocompletion script for the specified shell [nodus completion bash ](/docs/reference/cli/nodus_completion_bash/)Generate the autocompletion script for bash [nodus completion fish ](/docs/reference/cli/nodus_completion_fish/)Generate the autocompletion script for fish [nodus completion powershell ](/docs/reference/cli/nodus_completion_powershell/)Generate the autocompletion script for powershell [nodus completion zsh ](/docs/reference/cli/nodus_completion_zsh/)Generate the autocompletion script for zsh [nodus config ](/docs/reference/cli/nodus_config/)Show and change contexts in \~/.nodus/config [nodus config current-context ](/docs/reference/cli/nodus_config_current-context/)Print the current context [nodus config get-contexts ](/docs/reference/cli/nodus_config_get-contexts/)List contexts [nodus config set-context ](/docs/reference/cli/nodus_config_set-context/)Change a context's project or server (NAME defaults to the current context) [nodus config use-context ](/docs/reference/cli/nodus_config_use-context/)Switch the current context [nodus config view ](/docs/reference/cli/nodus_config_view/)Print the config (keys stay in the keychain) [nodus convert ](/docs/reference/cli/nodus_convert/)Turn a 0.x nodus.toml or training file into manifests for nodus apply [nodus cp ](/docs/reference/cli/nodus_cp/)Copy files to and from a running Sandbox, Workspace or Job, or download a Job output [nodus create ](/docs/reference/cli/nodus_create/)Create resources from manifests or with a typed generator [nodus create agentgroup ](/docs/reference/cli/nodus_create_agentgroup/)Create a agentgroup [nodus create agentrun ](/docs/reference/cli/nodus_create_agentrun/)Create a agentrun [nodus create apikey ](/docs/reference/cli/nodus_create_apikey/)Create an API key (printed once) [nodus create budget ](/docs/reference/cli/nodus_create_budget/)Create a budget [nodus create enrollmenttoken ](/docs/reference/cli/nodus_create_enrollmenttoken/)Create a token that enrolls one host into a Pool (installer printed once) [nodus create inferenceendpoint ](/docs/reference/cli/nodus_create_inferenceendpoint/)Create a inferenceendpoint [nodus create invite ](/docs/reference/cli/nodus_create_invite/)Invite someone to the org by email [nodus create sandbox ](/docs/reference/cli/nodus_create_sandbox/)Create a sandbox [nodus create secret ](/docs/reference/cli/nodus_create_secret/)Create a secret [nodus create sshkey ](/docs/reference/cli/nodus_create_sshkey/)Add an SSH public key, or generate a key pair and add it [nodus create token ](/docs/reference/cli/nodus_create_token/)Mint a ServiceAccount API key (printed once) [nodus create volume ](/docs/reference/cli/nodus_create_volume/)Create a volume [nodus create workspace ](/docs/reference/cli/nodus_create_workspace/)Create a workspace [nodus delete ](/docs/reference/cli/nodus_delete/)Delete resources [nodus deploy ](/docs/reference/cli/nodus_deploy/)Deploy the nodus.App in APP.py as a persistent App [nodus describe ](/docs/reference/cli/nodus_describe/)Show everything about one resource: attempts, placement, checkpoint, cost, conditions and events [nodus diff ](/docs/reference/cli/nodus_diff/)Show what apply would change, from a server dry-run (exit 1 when there are differences) [nodus edit ](/docs/reference/cli/nodus_edit/)Edit a resource in $EDITOR and save it as a merge patch [nodus events ](/docs/reference/cli/nodus_events/)List events in the project, or for one object, and watch for new ones [nodus exec ](/docs/reference/cli/nodus_exec/)Run a command in a running Job, Sandbox or Workspace (recorded as a Process) [nodus explain ](/docs/reference/cli/nodus_explain/)Show the documentation of a kind or one of its fields [nodus get ](/docs/reference/cli/nodus_get/)List or show resources [nodus get transactions ](/docs/reference/cli/nodus_get_transactions/)List postings on the org's wallet with the balance after each [nodus get usage ](/docs/reference/cli/nodus_get_usage/)Show metered usage, itemized or summed by project, kind, member, label, meter, day, segment, offering or rank [nodus inference ](/docs/reference/cli/nodus_inference/)Call hosted models: list them, send a chat completion and read a request's receipt [nodus inference chat ](/docs/reference/cli/nodus_inference_chat/)Send one chat completion and print the answer; the request id and charge go to stderr [nodus inference models ](/docs/reference/cli/nodus_inference_models/)List the models you can call now [nodus inference receipt ](/docs/reference/cli/nodus_inference_receipt/)Show what one inference request was charged (receipts are kept 30 days) [nodus init ](/docs/reference/cli/nodus_init/)Scaffold a manifest or app.py in the current directory [nodus label ](/docs/reference/cli/nodus_label/)Set or remove labels [nodus login ](/docs/reference/cli/nodus_login/)Sign in through the console and store one API key per chosen org [nodus logout ](/docs/reference/cli/nodus_logout/)Revoke the context's CLI key and remove the context [nodus logs ](/docs/reference/cli/nodus_logs/)Print or follow a resource's output [nodus mcp ](/docs/reference/cli/nodus_mcp/)Serve Nodus tools to a local MCP client over stdio [nodus mcp install ](/docs/reference/cli/nodus_mcp_install/)Add the Nodus MCP server to Claude Code, Cursor or Codex [nodus open ](/docs/reference/cli/nodus_open/)Open a resource's console page, or a port of its container, in the browser [nodus patch ](/docs/reference/cli/nodus_patch/)Update fields of a resource with a merge or JSON patch [nodus pool ](/docs/reference/cli/nodus_pool/)Pool utilization, forecasts, recommendations and actions [nodus pool approve ](/docs/reference/cli/nodus_pool_approve/)Approve a proposed PoolAction [nodus pool forecast ](/docs/reference/cli/nodus_pool_forecast/)Show a pool's demand forecast with p50 and p90 bands [nodus pool pause ](/docs/reference/cli/nodus_pool_pause/)Pause every automated action in a pool [nodus pool recommendations ](/docs/reference/cli/nodus_pool_recommendations/)List a pool's recommendations [nodus pool reject ](/docs/reference/cli/nodus_pool_reject/)Reject a proposed PoolAction [nodus pool resume ](/docs/reference/cli/nodus_pool_resume/)Resume every automated action in a pool [nodus pool revert ](/docs/reference/cli/nodus_pool_revert/)Undo an executed PoolAction where the action allows it [nodus pool utilization ](/docs/reference/cli/nodus_pool_utilization/)Show a pool's GPU utilization and device-hour ledger [nodus port-forward ](/docs/reference/cli/nodus_port-forward/)Forward local ports to a Job, Sandbox or Workspace [nodus request ](/docs/reference/cli/nodus_request/)Request an action: restart, refresh-secrets, reload-sinks, verify, resend, start, reimport, approve-burst, drain [nodus resume ](/docs/reference/cli/nodus_resume/)Resume a suspended run [nodus rollout ](/docs/reference/cli/nodus_rollout/)Roll out changes [nodus rollout restart ](/docs/reference/cli/nodus_rollout_restart/)Restart without a spec change (alias of \`request restart\`) [nodus run ](/docs/reference/cli/nodus_run/)Run a command as a Job: estimate, phases, logs and the final cost [nodus secret ](/docs/reference/cli/nodus_secret/)Create Secrets from literals, env files or registry logins [nodus secret create ](/docs/reference/cli/nodus_secret_create/)Create a Secret [nodus serve ](/docs/reference/cli/nodus_serve/)Run the App in APP.py and redeploy it whenever a file changes [nodus shell ](/docs/reference/cli/nodus_shell/)Open an interactive shell in a Sandbox, creating it if it is not there [nodus ssh ](/docs/reference/cli/nodus_ssh/)Connect to a Workspace over SSH [nodus start ](/docs/reference/cli/nodus_start/)Start a stopped Sandbox, Workspace, Function, Agent or InferenceEndpoint [nodus stop ](/docs/reference/cli/nodus_stop/)Stop a Sandbox, Workspace, Function, Agent or InferenceEndpoint [nodus suspend ](/docs/reference/cli/nodus_suspend/)Suspend a Job, Pipeline, Sweep, TrainingJob or AgentRun (checkpointing first) [nodus top ](/docs/reference/cli/nodus_top/)Show GPU, CPU and memory use, the hourly rate and spend of running Jobs [nodus version ](/docs/reference/cli/nodus_version/)Print the client version [nodus volume ](/docs/reference/cli/nodus_volume/)Put, get, list and remove the files of a Volume [nodus volume clear ](/docs/reference/cli/nodus_volume_clear/)Commit an empty revision; the Volume and its history stay [nodus volume get ](/docs/reference/cli/nodus_volume_get/)Download a file or directory from a Volume [nodus volume ls ](/docs/reference/cli/nodus_volume_ls/)List a directory of a Volume [nodus volume put ](/docs/reference/cli/nodus_volume_put/)Add a local file or directory to a Volume as a new revision [nodus volume rm ](/docs/reference/cli/nodus_volume_rm/)Remove a file or directory from a Volume as a new revision [nodus wait ](/docs/reference/cli/nodus_wait/)Wait until a resource meets a condition [nodus whoami ](/docs/reference/cli/nodus_whoami/)Show the signed-in principal, org, role, project, scopes and available credit # nodus > Run compute Jobs on Nodus ## nodus [Section titled “nodus”](#nodus) Run compute Jobs on Nodus ### Synopsis [Section titled “Synopsis”](#synopsis) nodus submits and manages Nodus resources: Jobs, Sandboxes, Functions, Agents and more. ### Options [Section titled “Options”](#options) ```plaintext --context string Context from the config file to use -h, --help help for nodus --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus agentrun](/docs/reference/cli/nodus_agentrun/) - Message an AgentRun and read its steps and answer * [nodus annotate](/docs/reference/cli/nodus_annotate/) - Set or remove annotations * [nodus api-resources](/docs/reference/cli/nodus_api-resources/) - List the kinds the server offers * [nodus apply](/docs/reference/cli/nodus_apply/) - Create or update resources from manifests (client-side three-way merge) * [nodus attach](/docs/reference/cli/nodus_attach/) - Attach to a running Process of a Job, Sandbox or Workspace * [nodus auth](/docs/reference/cli/nodus_auth/) - Credentials for other tools, and permission checks * [nodus billing](/docs/reference/cli/nodus_billing/) - Show the org’s balance; top up, redeem a code and manage auto-recharge * [nodus cancel](/docs/reference/cli/nodus_cancel/) - Cancel a run, FunctionCall or Process * [nodus cloud](/docs/reference/cli/nodus_cloud/) - Inspect connected cloud accounts * [nodus completion](/docs/reference/cli/nodus_completion/) - Generate the autocompletion script for the specified shell * [nodus config](/docs/reference/cli/nodus_config/) - Show and change contexts in \~/.nodus/config * [nodus convert](/docs/reference/cli/nodus_convert/) - Turn a 0.x nodus.toml or training file into manifests for nodus apply * [nodus cp](/docs/reference/cli/nodus_cp/) - Copy files to and from a running Sandbox, Workspace or Job, or download a Job output * [nodus create](/docs/reference/cli/nodus_create/) - Create resources from manifests or with a typed generator * [nodus delete](/docs/reference/cli/nodus_delete/) - Delete resources * [nodus deploy](/docs/reference/cli/nodus_deploy/) - Deploy the nodus.App in APP.py as a persistent App * [nodus describe](/docs/reference/cli/nodus_describe/) - Show everything about one resource: attempts, placement, checkpoint, cost, conditions and events * [nodus diff](/docs/reference/cli/nodus_diff/) - Show what apply would change, from a server dry-run (exit 1 when there are differences) * [nodus edit](/docs/reference/cli/nodus_edit/) - Edit a resource in $EDITOR and save it as a merge patch * [nodus events](/docs/reference/cli/nodus_events/) - List events in the project, or for one object, and watch for new ones * [nodus exec](/docs/reference/cli/nodus_exec/) - Run a command in a running Job, Sandbox or Workspace (recorded as a Process) * [nodus explain](/docs/reference/cli/nodus_explain/) - Show the documentation of a kind or one of its fields * [nodus get](/docs/reference/cli/nodus_get/) - List or show resources * [nodus inference](/docs/reference/cli/nodus_inference/) - Call hosted models: list them, send a chat completion and read a request’s receipt * [nodus init](/docs/reference/cli/nodus_init/) - Scaffold a manifest or app.py in the current directory * [nodus label](/docs/reference/cli/nodus_label/) - Set or remove labels * [nodus login](/docs/reference/cli/nodus_login/) - Sign in through the console and store one API key per chosen org * [nodus logout](/docs/reference/cli/nodus_logout/) - Revoke the context’s CLI key and remove the context * [nodus logs](/docs/reference/cli/nodus_logs/) - Print or follow a resource’s output * [nodus mcp](/docs/reference/cli/nodus_mcp/) - Serve Nodus tools to a local MCP client over stdio * [nodus open](/docs/reference/cli/nodus_open/) - Open a resource’s console page, or a port of its container, in the browser * [nodus patch](/docs/reference/cli/nodus_patch/) - Update fields of a resource with a merge or JSON patch * [nodus pool](/docs/reference/cli/nodus_pool/) - Pool utilization, forecasts, recommendations and actions * [nodus port-forward](/docs/reference/cli/nodus_port-forward/) - Forward local ports to a Job, Sandbox or Workspace * [nodus request](/docs/reference/cli/nodus_request/) - Request an action: restart, refresh-secrets, reload-sinks, verify, resend, start, reimport, approve-burst, drain * [nodus resume](/docs/reference/cli/nodus_resume/) - Resume a suspended run * [nodus rollout](/docs/reference/cli/nodus_rollout/) - Roll out changes * [nodus run](/docs/reference/cli/nodus_run/) - Run a command as a Job: estimate, phases, logs and the final cost * [nodus secret](/docs/reference/cli/nodus_secret/) - Create Secrets from literals, env files or registry logins * [nodus serve](/docs/reference/cli/nodus_serve/) - Run the App in APP.py and redeploy it whenever a file changes * [nodus shell](/docs/reference/cli/nodus_shell/) - Open an interactive shell in a Sandbox, creating it if it is not there * [nodus ssh](/docs/reference/cli/nodus_ssh/) - Connect to a Workspace over SSH * [nodus start](/docs/reference/cli/nodus_start/) - Start a stopped Sandbox, Workspace, Function, Agent or InferenceEndpoint * [nodus stop](/docs/reference/cli/nodus_stop/) - Stop a Sandbox, Workspace, Function, Agent or InferenceEndpoint * [nodus suspend](/docs/reference/cli/nodus_suspend/) - Suspend a Job, Pipeline, Sweep, TrainingJob or AgentRun (checkpointing first) * [nodus top](/docs/reference/cli/nodus_top/) - Show GPU, CPU and memory use, the hourly rate and spend of running Jobs * [nodus version](/docs/reference/cli/nodus_version/) - Print the client version * [nodus volume](/docs/reference/cli/nodus_volume/) - Put, get, list and remove the files of a Volume * [nodus wait](/docs/reference/cli/nodus_wait/) - Wait until a resource meets a condition * [nodus whoami](/docs/reference/cli/nodus_whoami/) - Show the signed-in principal, org, role, project, scopes and available credit # nodus agentrun > Message an AgentRun and read its steps and answer ## nodus agentrun [Section titled “nodus agentrun”](#nodus-agentrun) Message an AgentRun and read its steps and answer ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for agentrun ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus * [nodus agentrun answer](/docs/reference/cli/nodus_agentrun_answer/) - Print an AgentRun’s last answer in full * [nodus agentrun send](/docs/reference/cli/nodus_agentrun_send/) - Send a message to an AgentRun, which takes it as its next prompt * [nodus agentrun steps](/docs/reference/cli/nodus_agentrun_steps/) - List an AgentRun’s journal steps: what it did, in order # nodus agentrun answer > Print an AgentRun's last answer in full ## nodus agentrun answer [Section titled “nodus agentrun answer”](#nodus-agentrun-answer) Print an AgentRun’s last answer in full ```plaintext nodus agentrun answer NAME [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus agentrun answer triage ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for answer ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus agentrun](/docs/reference/cli/nodus_agentrun/) - Message an AgentRun and read its steps and answer # nodus agentrun send > Send a message to an AgentRun, which takes it as its next prompt ## nodus agentrun send [Section titled “nodus agentrun send”](#nodus-agentrun-send) Send a message to an AgentRun, which takes it as its next prompt ### Synopsis [Section titled “Synopsis”](#synopsis) Send a message to an AgentRun. A run created with spec.keepAlive waits for messages after each turn and wakes for this one; a running run reads it after its current turn. A finished run refuses it. ```plaintext nodus agentrun send NAME MESSAGE [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus agentrun send triage "Now check the closed ones too" nodus agentrun send triage "Retry" --key retry-1 ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for send --key string Message key: sending the same key again delivers the message once -o, --output string Output format: json prints the receipt ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus agentrun](/docs/reference/cli/nodus_agentrun/) - Message an AgentRun and read its steps and answer # nodus agentrun steps > List an AgentRun's journal steps: what it did, in order ## nodus agentrun steps [Section titled “nodus agentrun steps”](#nodus-agentrun-steps) List an AgentRun’s journal steps: what it did, in order ```plaintext nodus agentrun steps NAME [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus agentrun steps triage ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for steps -o, --output string Output format: json ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus agentrun](/docs/reference/cli/nodus_agentrun/) - Message an AgentRun and read its steps and answer # nodus annotate > Set or remove annotations ## nodus annotate [Section titled “nodus annotate”](#nodus-annotate) Set or remove annotations ```plaintext nodus annotate KIND/NAME KEY=VALUE... [KEY-] [flags] ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for annotate ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus annotate job/train note=hi ``` # nodus api-resources > List the kinds the server offers ## nodus api-resources [Section titled “nodus api-resources”](#nodus-api-resources) List the kinds the server offers ```plaintext nodus api-resources [flags] ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for api-resources ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus api-resources ``` # nodus apply > Create or update resources from manifests (client-side three-way merge) ## nodus apply [Section titled “nodus apply”](#nodus-apply) Create or update resources from manifests (client-side three-way merge) ```plaintext nodus apply -f FILE|DIR|- [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus apply -f job.yaml nodus apply -f job.yaml --dry-run=server -o estimate nodus apply -f manifests/ --prune -l app=trainer ``` ### Options [Section titled “Options”](#options) ```plaintext --dry-run string[="server"] server returns the estimate without creating; none, server or client (default "none") -f, --filename strings Manifest files or directories; - reads stdin -h, --help help for apply -o, --output string Output format: name, json, yaml, or estimate (with --dry-run=server) --prune Delete objects matching -l that the manifests no longer contain -R, --recursive Walk directories recursively -l, --selector string Label selector for --prune ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus apply -f keep.yaml --dry-run=client nodus apply -f keep.yaml -f old.yaml -f other.yaml nodus apply -f keep.yaml --prune -l app=trainer --dry-run=server nodus apply -f keep.yaml --prune -l app=trainer --dry-run=client -o name nodus apply -f keep.yaml --prune -l app=trainer nodus apply -f jobs.yaml nodus apply -f job.yaml nodus apply -f job-v2.yaml nodus apply -f job.yaml --dry-run=server -o estimate nodus apply -f bad.yaml nodus apply -f dev.yaml nodus apply -f pool.yaml ``` # nodus attach > Attach to a running Process of a Job, Sandbox or Workspace ## nodus attach [Section titled “nodus attach”](#nodus-attach) Attach to a running Process of a Job, Sandbox or Workspace ```plaintext nodus attach KIND/NAME --process N [-it] [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus attach sb/dev --process 12 -it nodus attach job/train --process 3 --rank 1 ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for attach --index int Job index (Indexed Jobs) (default -1) --process nodus get processes Process sequence number (as nodus get processes shows) or full name --rank string Gang member rank (default 0) -i, --stdin Pass stdin to the process -t, --tty The process has a terminal ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus attach sb/dev --process 1 ``` # nodus auth > Credentials for other tools, and permission checks ## nodus auth [Section titled “nodus auth”](#nodus-auth) Credentials for other tools, and permission checks ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for auth ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus * [nodus auth can-i](/docs/reference/cli/nodus_auth_can-i/) - Check whether the current credential may perform an action * [nodus auth token](/docs/reference/cli/nodus_auth_token/) - Print an ExecCredential for kubectl and client-go (the kubeconfig exec plugin) # nodus auth can-i > Check whether the current credential may perform an action ## nodus auth can-i [Section titled “nodus auth can-i”](#nodus-auth-can-i) Check whether the current credential may perform an action ```plaintext nodus auth can-i VERB RESOURCE [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus auth can-i create jobs nodus auth can-i exec sandboxes -p research ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for can-i ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus auth](/docs/reference/cli/nodus_auth/) - Credentials for other tools, and permission checks ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus auth can-i create jobs nodus auth can-i get jobs nodus auth can-i exec sandboxes ``` # nodus auth token > Print an ExecCredential for kubectl and client-go (the kubeconfig exec plugin) ## nodus auth token [Section titled “nodus auth token”](#nodus-auth-token) Print an ExecCredential for kubectl and client-go (the kubeconfig exec plugin) ```plaintext nodus auth token [flags] ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for token ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus auth](/docs/reference/cli/nodus_auth/) - Credentials for other tools, and permission checks ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus auth token ``` # nodus billing > Show the org's balance; top up, redeem a code and manage auto-recharge ## nodus billing [Section titled “nodus billing”](#nodus-billing) Show the org’s balance; top up, redeem a code and manage auto-recharge ### Synopsis [Section titled “Synopsis”](#synopsis) Without a subcommand, billing summarizes the org’s BillingAccount: what is available to spend, what open holds reserve, spend this month against the fullest budget, and auto-recharge. Billing changes need the billing:write scope (Owners and Admins). ```plaintext nodus billing [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus billing nodus billing top-up 20 nodus billing redeem LAUNCH25 nodus billing auto-recharge --threshold 10 --amount 50 ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for billing ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus * [nodus billing auto-recharge](/docs/reference/cli/nodus_billing_auto-recharge/) - Top up automatically when available credit falls below a threshold * [nodus billing deal](/docs/reference/cli/nodus_billing_deal/) - View partner offers or request an eligibility review * [nodus billing portal](/docs/reference/cli/nodus_billing_portal/) - Open the Stripe Customer Portal: cards, invoices and billing details * [nodus billing receipts](/docs/reference/cli/nodus_billing_receipts/) - List top-ups with their Stripe receipt links * [nodus billing redeem](/docs/reference/cli/nodus_billing_redeem/) - Redeem a promo code for credit * [nodus billing top-up](/docs/reference/cli/nodus_billing_top-up/) - Buy prepaid credit with Stripe Checkout ($5–$1,000) * [nodus billing usage](/docs/reference/cli/nodus_billing_usage/) - Show metered usage, itemized or summed by project, kind, member, label, meter, day, segment, offering or rank # nodus billing auto-recharge > Top up automatically when available credit falls below a threshold ## nodus billing auto-recharge [Section titled “nodus billing auto-recharge”](#nodus-billing-auto-recharge) Top up automatically when available credit falls below a threshold ### Synopsis [Section titled “Synopsis”](#synopsis) auto-recharge charges the saved card for –amount whenever available credit falls below –threshold. With no card on file it opens Stripe’s card setup first and turns auto-recharge on once the card is saved. Three failed charges in a row turn it off and notify the org’s Owners and Admins. ```plaintext nodus billing auto-recharge --threshold USD --amount USD | --off [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus billing auto-recharge --threshold 10 --amount 50 nodus billing auto-recharge --off ``` ### Options [Section titled “Options”](#options) ```plaintext --amount string Dollars to add each time ($5–$1,000) (default "20") -h, --help help for auto-recharge --no-open Print the card setup link instead of opening the browser --no-wait With no card on file, print the setup link and return without turning auto-recharge on --off Turn auto-recharge off --threshold string Recharge when available credit falls below this many dollars (at least 5) (default "10") ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus billing](/docs/reference/cli/nodus_billing/) - Show the org’s balance; top up, redeem a code and manage auto-recharge # nodus billing deal > View partner offers or request an eligibility review ## nodus billing deal [Section titled “nodus billing deal”](#nodus-billing-deal) View partner offers or request an eligibility review ```plaintext nodus billing deal [CODE] [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus billing deal nodus billing deal YC --note 'Company and batch' ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for deal --note string Company or batch information for eligibility review ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus billing](/docs/reference/cli/nodus_billing/) - Show the org’s balance; top up, redeem a code and manage auto-recharge # nodus billing portal > Open the Stripe Customer Portal: cards, invoices and billing details ## nodus billing portal [Section titled “nodus billing portal”](#nodus-billing-portal) Open the Stripe Customer Portal: cards, invoices and billing details ```plaintext nodus billing portal [flags] ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for portal --print Print the link instead of opening the browser --setup Add a card without buying credits ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus billing](/docs/reference/cli/nodus_billing/) - Show the org’s balance; top up, redeem a code and manage auto-recharge # nodus billing receipts > List top-ups with their Stripe receipt links ## nodus billing receipts [Section titled “nodus billing receipts”](#nodus-billing-receipts) List top-ups with their Stripe receipt links ```plaintext nodus billing receipts [flags] ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for receipts ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus billing](/docs/reference/cli/nodus_billing/) - Show the org’s balance; top up, redeem a code and manage auto-recharge # nodus billing redeem > Redeem a promo code for credit ## nodus billing redeem [Section titled “nodus billing redeem”](#nodus-billing-redeem) Redeem a promo code for credit ### Synopsis [Section titled “Synopsis”](#synopsis) redeem adds the code’s credit to the org as a CreditGrant. Redeeming the same code again changes nothing. ```plaintext nodus billing redeem CODE [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus billing redeem LAUNCH25 ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for redeem ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus billing](/docs/reference/cli/nodus_billing/) - Show the org’s balance; top up, redeem a code and manage auto-recharge # nodus billing top-up > Buy prepaid credit with Stripe Checkout ($5–$1,000) ## nodus billing top-up [Section titled “nodus billing top-up”](#nodus-billing-top-up) Buy prepaid credit with Stripe Checkout ($5–$1,000) ### Synopsis [Section titled “Synopsis”](#synopsis) top-up creates a TopUp, opens its Stripe Checkout page and waits until the payment is credited. Top-ups settle any arrears before they credit the wallet. Stripe emails the receipt to the billing email. ```plaintext nodus billing top-up USD [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus billing top-up 20 nodus billing top-up 50 --save-card nodus billing top-up 20 --no-open --no-wait ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for top-up --no-open Print the Checkout link instead of opening the browser --no-wait Return once the Checkout link exists, without waiting for the payment --save-card Save the card for auto-recharge ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus billing](/docs/reference/cli/nodus_billing/) - Show the org’s balance; top up, redeem a code and manage auto-recharge # nodus billing usage > Show metered usage, itemized or summed by project, kind, member, label, meter, day, segment, offering or rank ## nodus billing usage [Section titled “nodus billing usage”](#nodus-billing-usage) Show metered usage, itemized or summed by project, kind, member, label, meter, day, segment, offering or rank ### Synopsis [Section titled “Synopsis”](#synopsis) usage lists UsageRecords, or with –group-by sums them into a UsageSummary. Group keys are project, kind, meter, day, segment (Boot, Running, Restore, Teardown), offering (the capacity it ran on), rank (a gang member), member (the person who started the work, with any of their keys) and label:. Every billed interval of your allocation is on a line, from its billing start to confirmed deletion. ```plaintext nodus billing usage [--group-by KEY,...] [--since 30d] [--until TIME] [-o csv|json|yaml|wide] [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus get usage --group-by member --since 30d nodus get usage --group-by project,label:team --since 30d nodus get usage --object job/ddp-2node --group-by rank,segment nodus get usage --group-by day,meter --since 7d -o csv > usage.csv ``` ### Options [Section titled “Options”](#options) ```plaintext --field-selector string Field selector, such as meter=ComputeSeconds --group-by strings Sum by these keys: project, kind, meter, day, segment, offering, rank, member, label: -h, --help help for usage --no-headers Omit table headers --object string Only the usage of one object, as KIND/NAME -o, --output string Output format: wide, csv, json or yaml --since string Start of the range: a duration back from now (30d, 12h) or a time --until string End of the range (default now) ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus billing](/docs/reference/cli/nodus_billing/) - Show the org’s balance; top up, redeem a code and manage auto-recharge # nodus cancel > Cancel a run, FunctionCall or Process ## nodus cancel [Section titled “nodus cancel”](#nodus-cancel) Cancel a run, FunctionCall or Process ```plaintext nodus cancel KIND/NAME [flags] ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for cancel ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus cancel job/train ``` # nodus cloud > Inspect connected cloud accounts ## nodus cloud [Section titled “nodus cloud”](#nodus-cloud) Inspect connected cloud accounts ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for cloud ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus * [nodus cloud inventory](/docs/reference/cli/nodus_cloud_inventory/) - List the instances, GPUs or spend a cloud account’s last sync saw # nodus cloud inventory > List the instances, GPUs or spend a cloud account's last sync saw ## nodus cloud inventory [Section titled “nodus cloud inventory”](#nodus-cloud-inventory) List the instances, GPUs or spend a cloud account’s last sync saw ```plaintext nodus cloud inventory ACCOUNT [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus cloud inventory aws-main nodus cloud inventory aws-main --resource gpus ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for inventory --json Print the response as JSON --resource string instances, gpus or spend (default instances) ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus cloud](/docs/reference/cli/nodus_cloud/) - Inspect connected cloud accounts ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus cloud inventory aws-main ``` # nodus completion > Generate the autocompletion script for the specified shell ## nodus completion [Section titled “nodus completion”](#nodus-completion) Generate the autocompletion script for the specified shell ### Synopsis [Section titled “Synopsis”](#synopsis) Generate the autocompletion script for nodus for the specified shell. See each sub-command’s help for details on how to use the generated script. ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for completion ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus * [nodus completion bash](/docs/reference/cli/nodus_completion_bash/) - Generate the autocompletion script for bash * [nodus completion fish](/docs/reference/cli/nodus_completion_fish/) - Generate the autocompletion script for fish * [nodus completion powershell](/docs/reference/cli/nodus_completion_powershell/) - Generate the autocompletion script for powershell * [nodus completion zsh](/docs/reference/cli/nodus_completion_zsh/) - Generate the autocompletion script for zsh # nodus completion bash > Generate the autocompletion script for bash ## nodus completion bash [Section titled “nodus completion bash”](#nodus-completion-bash) Generate the autocompletion script for bash ### Synopsis [Section titled “Synopsis”](#synopsis) Generate the autocompletion script for the bash shell. This script depends on the ‘bash-completion’ package. If it is not installed already, you can install it via your OS’s package manager. To load completions in your current shell session: ```plaintext source <(nodus completion bash) ``` To load completions for every new session, execute once: #### Linux: [Section titled “Linux:”](#linux) ```plaintext nodus completion bash > /etc/bash_completion.d/nodus ``` #### macOS: [Section titled “macOS:”](#macos) ```plaintext nodus completion bash > $(brew --prefix)/etc/bash_completion.d/nodus ``` You will need to start a new shell for this setup to take effect. ```plaintext nodus completion bash ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for bash --no-descriptions disable completion descriptions ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus completion](/docs/reference/cli/nodus_completion/) - Generate the autocompletion script for the specified shell # nodus completion fish > Generate the autocompletion script for fish ## nodus completion fish [Section titled “nodus completion fish”](#nodus-completion-fish) Generate the autocompletion script for fish ### Synopsis [Section titled “Synopsis”](#synopsis) Generate the autocompletion script for the fish shell. To load completions in your current shell session: ```plaintext nodus completion fish | source ``` To load completions for every new session, execute once: ```plaintext nodus completion fish > ~/.config/fish/completions/nodus.fish ``` You will need to start a new shell for this setup to take effect. ```plaintext nodus completion fish [flags] ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for fish --no-descriptions disable completion descriptions ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus completion](/docs/reference/cli/nodus_completion/) - Generate the autocompletion script for the specified shell # nodus completion powershell > Generate the autocompletion script for powershell ## nodus completion powershell [Section titled “nodus completion powershell”](#nodus-completion-powershell) Generate the autocompletion script for powershell ### Synopsis [Section titled “Synopsis”](#synopsis) Generate the autocompletion script for powershell. To load completions in your current shell session: ```plaintext nodus completion powershell | Out-String | Invoke-Expression ``` To load completions for every new session, add the output of the above command to your powershell profile. ```plaintext nodus completion powershell [flags] ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for powershell --no-descriptions disable completion descriptions ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus completion](/docs/reference/cli/nodus_completion/) - Generate the autocompletion script for the specified shell # nodus completion zsh > Generate the autocompletion script for zsh ## nodus completion zsh [Section titled “nodus completion zsh”](#nodus-completion-zsh) Generate the autocompletion script for zsh ### Synopsis [Section titled “Synopsis”](#synopsis) Generate the autocompletion script for the zsh shell. If shell completion is not already enabled in your environment you will need to enable it. You can execute the following once: ```plaintext echo "autoload -U compinit; compinit" >> ~/.zshrc ``` To load completions in your current shell session: ```plaintext source <(nodus completion zsh) ``` To load completions for every new session, execute once: #### Linux: [Section titled “Linux:”](#linux) ```plaintext nodus completion zsh > "${fpath[1]}/_nodus" ``` #### macOS: [Section titled “macOS:”](#macos) ```plaintext nodus completion zsh > $(brew --prefix)/share/zsh/site-functions/_nodus ``` You will need to start a new shell for this setup to take effect. ```plaintext nodus completion zsh [flags] ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for zsh --no-descriptions disable completion descriptions ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus completion](/docs/reference/cli/nodus_completion/) - Generate the autocompletion script for the specified shell ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus completion zsh ``` # nodus config > Show and change contexts in ~/.nodus/config ## nodus config [Section titled “nodus config”](#nodus-config) Show and change contexts in \~/.nodus/config ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for config ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus * [nodus config current-context](/docs/reference/cli/nodus_config_current-context/) - Print the current context * [nodus config get-contexts](/docs/reference/cli/nodus_config_get-contexts/) - List contexts * [nodus config set-context](/docs/reference/cli/nodus_config_set-context/) - Change a context’s project or server (NAME defaults to the current context) * [nodus config use-context](/docs/reference/cli/nodus_config_use-context/) - Switch the current context * [nodus config view](/docs/reference/cli/nodus_config_view/) - Print the config (keys stay in the keychain) # nodus config current-context > Print the current context ## nodus config current-context [Section titled “nodus config current-context”](#nodus-config-current-context) Print the current context ```plaintext nodus config current-context [flags] ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for current-context ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus config](/docs/reference/cli/nodus_config/) - Show and change contexts in \~/.nodus/config ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus config current-context ``` # nodus config get-contexts > List contexts ## nodus config get-contexts [Section titled “nodus config get-contexts”](#nodus-config-get-contexts) List contexts ```plaintext nodus config get-contexts [flags] ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for get-contexts ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus config](/docs/reference/cli/nodus_config/) - Show and change contexts in \~/.nodus/config ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus config get-contexts ``` # nodus config set-context > Change a context's project or server (NAME defaults to the current context) ## nodus config set-context [Section titled “nodus config set-context”](#nodus-config-set-context) Change a context’s project or server (NAME defaults to the current context) ```plaintext nodus config set-context NAME [--project P] [--server URL] [flags] ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for set-context --project string Project for this context --server string API server for this context ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus config](/docs/reference/cli/nodus_config/) - Show and change contexts in \~/.nodus/config ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus config set-context acme --project research ``` # nodus config use-context > Switch the current context ## nodus config use-context [Section titled “nodus config use-context”](#nodus-config-use-context) Switch the current context ```plaintext nodus config use-context NAME [flags] ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for use-context ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus config](/docs/reference/cli/nodus_config/) - Show and change contexts in \~/.nodus/config ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus config use-context nope ``` # nodus config view > Print the config (keys stay in the keychain) ## nodus config view [Section titled “nodus config view”](#nodus-config-view) Print the config (keys stay in the keychain) ```plaintext nodus config view [flags] ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for view ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus config](/docs/reference/cli/nodus_config/) - Show and change contexts in \~/.nodus/config # nodus convert > Turn a 0.x nodus.toml or training file into manifests for nodus apply ## nodus convert [Section titled “nodus convert”](#nodus-convert) Turn a 0.x nodus.toml or training file into manifests for nodus apply ### Synopsis [Section titled “Synopsis”](#synopsis) convert maps a 0.x nodus.toml or nodus.json onto a Job (a Pipeline when it has stages), and a nodus.training.json onto a TrainingRuntime and a TrainingJob. Every field that does not carry over is named on stderr; nothing is sent to the API. ```plaintext nodus convert nodus.toml | nodus.json | nodus.training.json [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus convert nodus.toml > job.yaml && nodus apply -f job.yaml nodus convert nodus.training.json --image ghcr.io/acme/trainer@sha256:... | nodus apply -f - ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for convert --image string Trainer image for a nodus.training.json runtime --name string metadata.name for the manifests (default: the file's directory name) -o, --output string Output format: yaml or json (default "yaml") ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus convert nodus.toml nodus convert --name 'Train Eval' nodus.toml nodus convert -o json payload/nodus.json nodus convert --image ghcr.io/acme/trainer@sha256:0123 training/nodus.training.json ``` # nodus cp > Copy files to and from a running Sandbox, Workspace or Job, or download a Job output ## nodus cp [Section titled “nodus cp”](#nodus-cp) Copy files to and from a running Sandbox, Workspace or Job, or download a Job output ### Synopsis [Section titled “Synopsis”](#synopsis) Either side is KIND/NAME:/absolute/path inside the running container; the other side is a local path. Directories copy recursively. KIND/NAME:outputs/OUTPUT downloads a declared output of a Job, verified against its SHA-256 and written atomically; KIND/NAME:outputs/ or outputs/PREFIX/ downloads every output under it into a directory, so a TrainingJob’s adapter comes down in one command. ```plaintext nodus cp SRC DST [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus cp ./data sb/dev:/workspace/data nodus cp ws/lab:/home/nodus/results ./results nodus cp job/train:outputs/model ./model nodus cp trainingjob/tune:outputs/ ./outputs nodus cp job/ddp:outputs/metrics . --rank 1 ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for cp --index int Job index (Indexed Jobs) (default -1) --rank string Gang member rank (default 0) ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus cp job/train:outputs/model model.bin nodus cp job/train:outputs/model out nodus cp want-model.bin sb/dev:/workspace/ nodus cp src sb/dev:/workspace/src nodus cp sb/dev:/workspace back nodus cp sb/dev:/workspace/src/a.bin single.bin nodus cp sb/dev:/workspace/missing gone.bin nodus cp want-model.bin sb/dev:/hostile/ok.bin nodus cp sb/dev:/hostile got ``` # nodus create > Create resources from manifests or with a typed generator ## nodus create [Section titled “nodus create”](#nodus-create) Create resources from manifests or with a typed generator ```plaintext nodus create -f FILE | create KIND NAME [flags] ``` ### Options [Section titled “Options”](#options) ```plaintext --dry-run string[="server"] server returns the estimate without creating; none, server or client (default "none") -f, --filename strings Manifest files or directories; - reads stdin -h, --help help for create -o, --output string Output format: name, json, yaml, or estimate (with --dry-run=server) -R, --recursive Walk directories recursively ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus * [nodus create agentgroup](/docs/reference/cli/nodus_create_agentgroup/) - Create a agentgroup * [nodus create agentrun](/docs/reference/cli/nodus_create_agentrun/) - Create a agentrun * [nodus create apikey](/docs/reference/cli/nodus_create_apikey/) - Create an API key (printed once) * [nodus create budget](/docs/reference/cli/nodus_create_budget/) - Create a budget * [nodus create enrollmenttoken](/docs/reference/cli/nodus_create_enrollmenttoken/) - Create a token that enrolls one host into a Pool (installer printed once) * [nodus create inferenceendpoint](/docs/reference/cli/nodus_create_inferenceendpoint/) - Create a inferenceendpoint * [nodus create invite](/docs/reference/cli/nodus_create_invite/) - Invite someone to the org by email * [nodus create sandbox](/docs/reference/cli/nodus_create_sandbox/) - Create a sandbox * [nodus create secret](/docs/reference/cli/nodus_create_secret/) - Create a secret * [nodus create sshkey](/docs/reference/cli/nodus_create_sshkey/) - Add an SSH public key, or generate a key pair and add it * [nodus create token](/docs/reference/cli/nodus_create_token/) - Mint a ServiceAccount API key (printed once) * [nodus create volume](/docs/reference/cli/nodus_create_volume/) - Create a volume * [nodus create workspace](/docs/reference/cli/nodus_create_workspace/) - Create a workspace ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus create -f generated.yaml -o name nodus create -f stray.yaml ``` # nodus create agentgroup > Create a agentgroup ## nodus create agentgroup [Section titled “nodus create agentgroup”](#nodus-create-agentgroup) Create a agentgroup ```plaintext nodus create agentgroup NAME [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus create agentgroup triage --agent researcher --max-active 8 --max-cost 30 ``` ### Options [Section titled “Options”](#options) ```plaintext --agent string Agent every run of the group runs --dry-run string[="server"] server returns the estimate without creating; client prints the manifest (default "none") -h, --help help for agentgroup --max-active int Runs that run at once --max-cost string Spend limit in USD shared by every run --max-pending int Unfinished runs the group admits -o, --output string Output format: name, json, yaml, or estimate (with --dry-run=server) ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus create](/docs/reference/cli/nodus_create/) - Create resources from manifests or with a typed generator # nodus create agentrun > Create a agentrun ## nodus create agentrun [Section titled “nodus create agentrun”](#nodus-create-agentrun) Create a agentrun ```plaintext nodus create agentrun NAME [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus create agentrun triage --agent researcher --prompt "Summarize the open issues" ``` ### Options [Section titled “Options”](#options) ```plaintext --agent string Agent to run: one in the project, or a template of project nodus --dry-run string[="server"] server returns the estimate without creating; client prints the manifest (default "none") -h, --help help for agentrun -o, --output string Output format: name, json, yaml, or estimate (with --dry-run=server) --prompt string The first prompt ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus create](/docs/reference/cli/nodus_create/) - Create resources from manifests or with a typed generator # nodus create apikey > Create an API key (printed once) ## nodus create apikey [Section titled “nodus create apikey”](#nodus-create-apikey) Create an API key (printed once) ```plaintext nodus create apikey NAME [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus create apikey ci --scopes jobs:write,volumes:read --projects research --expires 720h ``` ### Options [Section titled “Options”](#options) ```plaintext --description string What the key is for (at most 256 characters) --dry-run string[="server"] server returns the estimate without creating; client prints the manifest (default "none") --expires string Lifetime of the key, such as 720h (default: it does not expire) -h, --help help for apikey -o, --output string Output format: name, json, yaml, or estimate (with --dry-run=server) --projects strings Projects the key may reach, comma-separated or repeated (default: all) --scopes strings Scopes such as jobs:write, comma-separated or repeated (default: whatever your role allows) --service-account string Bind the key to this ServiceAccount of the project (-p) ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus create](/docs/reference/cli/nodus_create/) - Create resources from manifests or with a typed generator ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus create apikey ci --scopes jobs:write,volumes:read --projects research --expires 720h nodus create apikey ci2 -o name nodus create apikey bot --service-account deploy --dry-run=client ``` # nodus create budget > Create a budget ## nodus create budget [Section titled “nodus create budget”](#nodus-create-budget) Create a budget ```plaintext nodus create budget NAME [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus create budget research --limit 2000 --period Monthly --scope-project research nodus create budget ada-monthly --limit 300 --period Monthly --scope-member ada@example.com ``` ### Options [Section titled “Options”](#options) ```plaintext --action string Block or Notify --dry-run string[="server"] server returns the estimate without creating; client prints the manifest (default "none") -h, --help help for budget --limit string Limit in USD -o, --output string Output format: name, json, yaml, or estimate (with --dry-run=server) --period string Monthly or Total --scope-member string Limit spend of one member, by email, across every key they use --scope-project string Limit spend of one project ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus create](/docs/reference/cli/nodus_create/) - Create resources from manifests or with a typed generator ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus create budget research --limit 2000 --period Monthly --scope-project research ``` # nodus create enrollmenttoken > Create a token that enrolls one host into a Pool (installer printed once) ## nodus create enrollmenttoken [Section titled “nodus create enrollmenttoken”](#nodus-create-enrollmenttoken) Create a token that enrolls one host into a Pool (installer printed once) ```plaintext nodus create enrollmenttoken [NAME] [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus create enrollmenttoken --pool lab-a100 --ttl 2h ``` ### Options [Section titled “Options”](#options) ```plaintext --dry-run string[="server"] server returns the estimate without creating; client prints the manifest (default "none") -h, --help help for enrollmenttoken --label stringArray KEY=VALUE label copied onto the Node when it enrolls (repeatable) -o, --output string Output format: name, json, yaml, or estimate (with --dry-run=server) --pool string Pool the host joins --ttl string How long the token works, such as 2h (at most 24h; default 24h) ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus create](/docs/reference/cli/nodus_create/) - Create resources from manifests or with a typed generator ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus create enrollmenttoken --pool lab --ttl 2h --label rack=a ``` # nodus create inferenceendpoint > Create a inferenceendpoint ## nodus create inferenceendpoint [Section titled “nodus create inferenceendpoint”](#nodus-create-inferenceendpoint) Create a inferenceendpoint ```plaintext nodus create inferenceendpoint NAME [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus create inferenceendpoint support-bot --model nodus/gpt-oss-120b --rpm 120 --max-cost 5 ``` ### Options [Section titled “Options”](#options) ```plaintext --allowed-key stringArray API key name allowed to call it (repeatable) --dry-run string[="server"] server returns the estimate without creating; client prints the manifest (default "none") -h, --help help for inferenceendpoint --max-concurrent int Open requests --max-cost string Spend limit in USD --model string Catalog model, as 'nodus inference models' lists them -o, --output string Output format: name, json, yaml, or estimate (with --dry-run=server) --rpm int Requests per minute --tpm int Tokens per minute ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus create](/docs/reference/cli/nodus_create/) - Create resources from manifests or with a typed generator # nodus create invite > Invite someone to the org by email ## nodus create invite [Section titled “nodus create invite”](#nodus-create-invite) Invite someone to the org by email ```plaintext nodus create invite --email EMAIL [--role ROLE] [--projects P,...] [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus create invite --email ada@example.com --role Admin nodus create invite --email lin@example.com --projects research,evals ``` ### Options [Section titled “Options”](#options) ```plaintext --email string The invitee's email address; they accept by signing in with it -h, --help help for invite --projects strings Projects the member may reach, comma-separated (default: all) --role string Owner, Admin, Member or Viewer (default "Member") ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus create](/docs/reference/cli/nodus_create/) - Create resources from manifests or with a typed generator # nodus create sandbox > Create a sandbox ## nodus create sandbox [Section titled “nodus create sandbox”](#nodus-create-sandbox) Create a sandbox ```plaintext nodus create sandbox NAME [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus create sandbox dev --cpu 2 --memory 4Gi --idle-timeout 10m ``` ### Options [Section titled “Options”](#options) ```plaintext --cpu string vCPU floor --disk string Disk floor, such as 50Gi --dry-run string[="server"] server returns the estimate without creating; client prints the manifest (default "none") --gpu string Accelerator TYPE[:COUNT], such as L4 or H100:8 -h, --help help for sandbox --idle-timeout string Stop after this idle time --image string Container image (default "nodus/agent-tools") --max-cost string Spend limit in USD --memory string Memory floor, such as 16Gi -o, --output string Output format: name, json, yaml, or estimate (with --dry-run=server) --secret stringArray Secret to mount (repeatable) ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus create](/docs/reference/cli/nodus_create/) - Create resources from manifests or with a typed generator ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus create sandbox dev --cpu 2 --secret gh-token --dry-run=client nodus create sandbox dev --cpu 2 ``` # nodus create secret > Create a secret ## nodus create secret [Section titled “nodus create secret”](#nodus-create-secret) Create a secret ```plaintext nodus create secret NAME [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus create secret hf --from-literal HF_TOKEN=hf_xxx ``` ### Options [Section titled “Options”](#options) ```plaintext --dry-run string[="server"] server returns the estimate without creating; client prints the manifest (default "none") --from-env-file stringArray File of KEY=VALUE lines --from-literal stringArray KEY=VALUE (repeatable) -h, --help help for secret -o, --output string Output format: name, json, yaml, or estimate (with --dry-run=server) --type string Opaque or Registry ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus create](/docs/reference/cli/nodus_create/) - Create resources from manifests or with a typed generator ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus create secret hf --from-literal HF_TOKEN=hf_test --from-env-file wandb.env -o jsonpath={.spec.stringData} ``` # nodus create sshkey > Add an SSH public key, or generate a key pair and add it ## nodus create sshkey [Section titled “nodus create sshkey”](#nodus-create-sshkey) Add an SSH public key, or generate a key pair and add it ```plaintext nodus create sshkey [NAME] [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus create sshkey laptop --from-file ~/.ssh/id_ed25519.pub nodus create sshkey --generate ``` ### Options [Section titled “Options”](#options) ```plaintext --display-name string A label for the console --dry-run string[="server"] server returns the estimate without creating; client prints the manifest (default "none") --from-file string OpenSSH public key file to add, or - for standard input --generate Create a new Ed25519 key pair here and add its public key -h, --help help for sshkey --key-file string With --generate: where to write the private key (the public key goes next to it, with .pub) (default "~/.ssh/id_ed25519") -o, --output string Output format: name, json, yaml, or estimate (with --dry-run=server) ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus create](/docs/reference/cli/nodus_create/) - Create resources from manifests or with a typed generator ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus create sshkey laptop --from-file laptop.pub nodus create sshkey fresh --generate --key-file id_fresh ``` # nodus create token > Mint a ServiceAccount API key (printed once) ## nodus create token [Section titled “nodus create token”](#nodus-create-token) Mint a ServiceAccount API key (printed once) ```plaintext nodus create token sa/NAME [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus create token sa/ci --duration 720h --scope jobs:write ``` ### Options [Section titled “Options”](#options) ```plaintext --duration duration Lifetime of the key, such as 720h (default the server's 90 days; at most 8760h) -h, --help help for token --scope stringArray Scope such as jobs:write (repeatable; default the ServiceAccount's) ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus create](/docs/reference/cli/nodus_create/) - Create resources from manifests or with a typed generator ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus create token sa/ci --duration 720h --scope jobs:write nodus create token sa/ci ``` # nodus create volume > Create a volume ## nodus create volume [Section titled “nodus create volume”](#nodus-create-volume) Create a volume ```plaintext nodus create volume NAME [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus create volume weights --size 200Gi --access-mode ReadOnlyMany ``` ### Options [Section titled “Options”](#options) ```plaintext --access-mode string ReadWriteOnce, ReadOnlyMany or ReadWriteMany --dry-run string[="server"] server returns the estimate without creating; client prints the manifest (default "none") -h, --help help for volume -o, --output string Output format: name, json, yaml, or estimate (with --dry-run=server) --size string Capacity, such as 100Gi ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus create](/docs/reference/cli/nodus_create/) - Create resources from manifests or with a typed generator # nodus create workspace > Create a workspace ## nodus create workspace [Section titled “nodus create workspace”](#nodus-create-workspace) Create a workspace ```plaintext nodus create workspace NAME [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus create workspace lab --gpu L4 --image nodus/workspace-pytorch-cuda ``` ### Options [Section titled “Options”](#options) ```plaintext --cpu string vCPU floor --disk string Disk floor, such as 50Gi --dry-run string[="server"] server returns the estimate without creating; client prints the manifest (default "none") --gpu string Accelerator TYPE[:COUNT], such as L4 or H100:8 -h, --help help for workspace --image string Container image --memory string Memory floor, such as 16Gi -o, --output string Output format: name, json, yaml, or estimate (with --dry-run=server) --secret stringArray Secret to mount (repeatable) ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus create](/docs/reference/cli/nodus_create/) - Create resources from manifests or with a typed generator # nodus delete > Delete resources ## nodus delete [Section titled “nodus delete”](#nodus-delete) Delete resources ```plaintext nodus delete KIND/NAME... | -f FILE [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus delete job/finetune --wait nodus delete -f job.yaml ``` ### Options [Section titled “Options”](#options) ```plaintext --cascade string background, foreground or orphan (default "background") -f, --filename strings Manifests naming the resources to delete -h, --help help for delete --ignore-not-found Succeed when a resource is already gone -l, --selector string Label selector --timeout duration How long --wait waits (default 10m0s) --wait Wait until the resources are gone (finalizers cleared) ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus delete job/train nodus delete job/train --ignore-not-found nodus delete job/train --wait nodus delete sshkey laptop ``` # nodus deploy > Deploy the nodus.App in APP.py as a persistent App ## nodus deploy [Section titled “nodus deploy”](#nodus-deploy) Deploy the nodus.App in APP.py as a persistent App ```plaintext nodus deploy APP.py [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus deploy app.py ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for deploy ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus deploy app.py ``` # nodus describe > Show everything about one resource: attempts, placement, checkpoint, cost, conditions and events ## nodus describe [Section titled “nodus describe”](#nodus-describe) Show everything about one resource: attempts, placement, checkpoint, cost, conditions and events ```plaintext nodus describe KIND/NAME [flags] ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for describe ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus describe job/train nodus describe job/late ``` # nodus diff > Show what apply would change, from a server dry-run (exit 1 when there are differences) ## nodus diff [Section titled “nodus diff”](#nodus-diff) Show what apply would change, from a server dry-run (exit 1 when there are differences) ```plaintext nodus diff -f FILE [flags] ``` ### Options [Section titled “Options”](#options) ```plaintext -f, --filename strings Manifest file, directory or - for stdin -h, --help help for diff -R, --recursive Read directories recursively ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus diff -f job.yaml nodus diff -f job-v2.yaml ``` # nodus edit > Edit a resource in $EDITOR and save it as a merge patch ## nodus edit [Section titled “nodus edit”](#nodus-edit) Edit a resource in $EDITOR and save it as a merge patch ```plaintext nodus edit KIND/NAME [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus edit job/finetune EDITOR=nano nodus edit budget/research-monthly ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for edit ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus edit job/train ``` # nodus events > List events in the project, or for one object, and watch for new ones ## nodus events [Section titled “nodus events”](#nodus-events) List events in the project, or for one object, and watch for new ones ```plaintext nodus events [--for KIND/NAME] [-w] [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus events nodus events --for job/finetune -w nodus events -w -o json ``` ### Options [Section titled “Options”](#options) ```plaintext --for string Only events about KIND/NAME -h, --help help for events -o, --output string Output format: json, yaml, name, jsonpath=..., go-template=... or custom-columns=... -w, --watch Print the events, then each new one as it is recorded ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus events --for job/train ``` # nodus exec > Run a command in a running Job, Sandbox or Workspace (recorded as a Process) ## nodus exec [Section titled “nodus exec”](#nodus-exec) Run a command in a running Job, Sandbox or Workspace (recorded as a Process) ```plaintext nodus exec [-it] KIND/NAME -- COMMAND [ARGS...] [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus exec -it job/finetune -- bash nodus exec sb/dev -- nvidia-smi ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for exec --index int Job index (Indexed Jobs) (default -1) --rank string Gang member rank (default 0) -i, --stdin Pass stdin to the command -t, --tty Allocate a terminal ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus exec sb/dev -- python -c 'raise SystemExit(3)' ``` # nodus explain > Show the documentation of a kind or one of its fields ## nodus explain [Section titled “nodus explain”](#nodus-explain) Show the documentation of a kind or one of its fields ```plaintext nodus explain KIND[.FIELD...] [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus explain jobs nodus explain job.spec.resources.gpu nodus explain sandbox.spec --recursive ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for explain -o, --output string Output format: plaintext or plaintext-openapiv2 (default "plaintext") --recursive Print every nested field ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus explain jobs nodus explain job.spec.maxCostUSD ``` # nodus get > List or show resources ## nodus get [Section titled “nodus get”](#nodus-get) List or show resources ```plaintext nodus get KIND[/NAME]... [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus get jobs nodus get job/finetune -o yaml nodus get jobs -l team=nlp -w nodus get gpus --gpu H100 --count 8 --region us ``` ### Options [Section titled “Options”](#options) ```plaintext -A, --all-projects List across every project --count int get gpus: accelerators per offering --field-selector string Field selector -f, --filename strings Files naming the resources to get --gpu string get gpus: accelerator family, such as H100 -h, --help help for get --interruptible get gpus: interruptible capacity only --mine Only resources I created --no-headers Omit table headers -o, --output string Output format: wide, json, yaml, name, jsonpath=..., go-template=..., custom-columns=HEADER:.path,... --region string get gpus: region class, such as us or eu -l, --selector string Label selector --show-labels Show labels as the last column -w, --watch Watch for changes after listing ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus * [nodus get transactions](/docs/reference/cli/nodus_get_transactions/) - List postings on the org’s wallet with the balance after each * [nodus get usage](/docs/reference/cli/nodus_get_usage/) - Show metered usage, itemized or summed by project, kind, member, label, meter, day, segment, offering or rank ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus get job/keep nodus get job/old -o name nodus get job/old nodus get jobs -o name nodus get jobs nodus get jobs -o wide nodus get job/train -o jsonpath={.metadata.labels.team} nodus get job train -o yaml nodus get jobs -l team=nlp nodus get job/train -o jsonpath={.spec.maxCostUSD} nodus get jobs -w -o custom-columns=NAME:.metadata.name,PHASE:.status.phase nodus get job/train ``` # nodus get transactions > List postings on the org's wallet with the balance after each ## nodus get transactions [Section titled “nodus get transactions”](#nodus-get-transactions) List postings on the org’s wallet with the balance after each ```plaintext nodus get transactions [NAME] [--since 7d] [--type TYPE] [-o csv|json|yaml|wide] [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus get transactions --since 7d nodus get transactions --type Capture -o csv ``` ### Options [Section titled “Options”](#options) ```plaintext --field-selector string Field selector, such as meter=ComputeSeconds -h, --help help for transactions --no-headers Omit table headers -o, --output string Output format: wide, csv, json or yaml --since string Start of the range: a duration back from now (30d, 12h) or a time --type string Only this type, such as TopUp, Capture, Grant or Refund --until string End of the range (default now) ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus get](/docs/reference/cli/nodus_get/) - List or show resources # nodus get usage > Show metered usage, itemized or summed by project, kind, member, label, meter, day, segment, offering or rank ## nodus get usage [Section titled “nodus get usage”](#nodus-get-usage) Show metered usage, itemized or summed by project, kind, member, label, meter, day, segment, offering or rank ### Synopsis [Section titled “Synopsis”](#synopsis) usage lists UsageRecords, or with –group-by sums them into a UsageSummary. Group keys are project, kind, meter, day, segment (Boot, Running, Restore, Teardown), offering (the capacity it ran on), rank (a gang member), member (the person who started the work, with any of their keys) and label:. Every billed interval of your allocation is on a line, from its billing start to confirmed deletion. ```plaintext nodus get usage [--group-by KEY,...] [--since 30d] [--until TIME] [-o csv|json|yaml|wide] [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus get usage --group-by member --since 30d nodus get usage --group-by project,label:team --since 30d nodus get usage --object job/ddp-2node --group-by rank,segment nodus get usage --group-by day,meter --since 7d -o csv > usage.csv ``` ### Options [Section titled “Options”](#options) ```plaintext --field-selector string Field selector, such as meter=ComputeSeconds --group-by strings Sum by these keys: project, kind, meter, day, segment, offering, rank, member, label: -h, --help help for usage --no-headers Omit table headers --object string Only the usage of one object, as KIND/NAME -o, --output string Output format: wide, csv, json or yaml --since string Start of the range: a duration back from now (30d, 12h) or a time --until string End of the range (default now) ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus get](/docs/reference/cli/nodus_get/) - List or show resources # nodus inference > Call hosted models: list them, send a chat completion and read a request's receipt ## nodus inference [Section titled “nodus inference”](#nodus-inference) Call hosted models: list them, send a chat completion and read a request’s receipt ### Synopsis [Section titled “Synopsis”](#synopsis) inference talks to the OpenAI-compatible data plane of the current context with its credential: the inference host of the API (inference. for api., else the API host itself), or –base-url. ### Options [Section titled “Options”](#options) ```plaintext --base-url string Inference origin (default: derived from the API URL) -h, --help help for inference ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus * [nodus inference chat](/docs/reference/cli/nodus_inference_chat/) - Send one chat completion and print the answer; the request id and charge go to stderr * [nodus inference models](/docs/reference/cli/nodus_inference_models/) - List the models you can call now * [nodus inference receipt](/docs/reference/cli/nodus_inference_receipt/) - Show what one inference request was charged (receipts are kept 30 days) # nodus inference chat > Send one chat completion and print the answer; the request id and charge go to stderr ## nodus inference chat [Section titled “nodus inference chat”](#nodus-inference-chat) Send one chat completion and print the answer; the request id and charge go to stderr ### Synopsis [Section titled “Synopsis”](#synopsis) chat sends PROMPT (or standard input, with - or no argument) as one user message and prints the answer. The request is held at its maximum cost (the input plus –max-tokens of output at the model’s rates) and charged for the tokens it used; the request id, the model that answered, the tokens and the charge are printed to stderr. ```plaintext nodus inference chat [PROMPT | -] [--model MODEL | --endpoint NAME] [--max-tokens N] [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus inference chat "What is the capital of France?" nodus inference chat --model nodus/gpt-oss-20b --max-tokens 64 "Say hello" nodus inference chat --endpoint support-bot < question.txt nodus inference chat -o json "Hi" | jq .usage ``` ### Options [Section titled “Options”](#options) ```plaintext --endpoint string Send through this InferenceEndpoint of the project instead of --model -h, --help help for chat --idempotency-key string Retry safely: a repeat with the same key and prompt is answered once and charged once --max-tokens int Most output tokens; the request's hold is sized from it (default 512) --model string Catalog model, as 'nodus inference models' lists them; nodus/indra (Indra) routes each request (default "nodus/indra") -o, --output string Output format: json or yaml (the whole completion) --system string System message ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --base-url string Inference origin (default: derived from the API URL) --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus inference](/docs/reference/cli/nodus_inference/) - Call hosted models: list them, send a chat completion and read a request’s receipt # nodus inference models > List the models you can call now ## nodus inference models [Section titled “nodus inference models”](#nodus-inference-models) List the models you can call now ```plaintext nodus inference models [-o json] [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus inference models ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for models -o, --output string Output format: json or yaml ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --base-url string Inference origin (default: derived from the API URL) --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus inference](/docs/reference/cli/nodus_inference/) - Call hosted models: list them, send a chat completion and read a request’s receipt # nodus inference receipt > Show what one inference request was charged (receipts are kept 30 days) ## nodus inference receipt [Section titled “nodus inference receipt”](#nodus-inference-receipt) Show what one inference request was charged (receipts are kept 30 days) ### Synopsis [Section titled “Synopsis”](#synopsis) receipt shows a request’s state (Running, Settled, Released or Unknown), its usage and its charge. Every response of the data plane names its request in the Nodus-Request-Id header. A Released or Unknown request is never charged. ```plaintext nodus inference receipt REQUEST_ID [-o json|yaml] [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus inference receipt ireq_01k6d2x7q9fvjt3y8m0c4r5n2e ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for receipt -o, --output string Output format: json or yaml ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --base-url string Inference origin (default: derived from the API URL) --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus inference](/docs/reference/cli/nodus_inference/) - Call hosted models: list them, send a chat completion and read a request’s receipt # nodus init > Scaffold a manifest or app.py in the current directory ## nodus init [Section titled “nodus init”](#nodus-init) Scaffold a manifest or app.py in the current directory ```plaintext nodus init [job|app|agent|sandbox] [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus init # job.yaml nodus init app # app.py for nodus run app.py ``` ### Options [Section titled “Options”](#options) ```plaintext --force Replace an existing file -h, --help help for init --name string Resource or App name (default: the directory name) ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus init nodus init app --name demo ``` # nodus label > Set or remove labels ## nodus label [Section titled “nodus label”](#nodus-label) Set or remove labels ```plaintext nodus label KIND/NAME KEY=VALUE... [KEY-] [flags] ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for label ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus label job/train tier=gold ``` # nodus login > Sign in through the console and store one API key per chosen org ## nodus login [Section titled “nodus login”](#nodus-login) Sign in through the console and store one API key per chosen org ### Synopsis [Section titled “Synopsis”](#synopsis) Opens the console, where you sign in (or sign up) and choose orgs; nodus stores one 90-day API key per org in the OS keychain and writes one context per org. Interactive sign-in also creates or reuses your local SSH key and configures access to all current and future Workspaces in those orgs. –device prints a code to enter on another device; –with-token reads an API key from stdin (CI). ```plaintext nodus login [--device | --with-token] [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus login nodus login --device echo "$NODUS_API_KEY" | nodus login --with-token ``` ### Options [Section titled “Options”](#options) ```plaintext --api-url string API server (default https://api.nodus-compute.ai) --console-url string Console origin (default: what the API server names) --device Sign in from another device with a one-time code (RFC 8628) -h, --help help for login --with-token Read an API key from stdin ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus login --with-token nodus login --device nodus login ``` # nodus logout > Revoke the context's CLI key and remove the context ## nodus logout [Section titled “nodus logout”](#nodus-logout) Revoke the context’s CLI key and remove the context ```plaintext nodus logout [flags] ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for logout ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus logout ``` # nodus logs > Print or follow a resource's output ## nodus logs [Section titled “nodus logs”](#nodus-logs) Print or follow a resource’s output ```plaintext nodus logs KIND/NAME [-f] [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus logs -f job/finetune nodus logs job/ddp --all-ranks nodus logs sb/dev --process 12 ``` ### Options [Section titled “Options”](#options) ```plaintext --all-indexes Every index of an Indexed Job --all-ranks Every gang member, each line prefixed with its rank --attempt int Attempt number (default: the current one) -f, --follow Stream new output -h, --help help for logs --index int Job index (Indexed Jobs) (default -1) --process string A Process of a Sandbox or Workspace, by sequence number --rank string Gang member rank (default 0) --since duration Only output newer than this, such as 10m --tail int Lines of backlog to show (default -1) --timestamps Prefix each line with its time ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus logs job/train nodus logs job/train --rank 1 nodus logs job/train --all-ranks nodus logs job/shards --all-indexes ``` # nodus mcp > Serve Nodus tools to a local MCP client over stdio ## nodus mcp [Section titled “nodus mcp”](#nodus-mcp) Serve Nodus tools to a local MCP client over stdio ### Synopsis [Section titled “Synopsis”](#synopsis) Runs the Nodus MCP server on standard input and output, for clients that launch it as a subprocess. It acts with this machine’s login (nodus login or NODUS_API_KEY), so it can do what that credential can, and its writes return a dry-run that the client’s user confirms before anything is created. On top of the hosted tools it can upload the current directory as source, download outputs and copy files to and from a Sandbox. To connect a client, run nodus mcp install CLIENT, or add the hosted server from . ```plaintext nodus mcp [flags] ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for mcp ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus * [nodus mcp install](/docs/reference/cli/nodus_mcp_install/) - Add the Nodus MCP server to Claude Code, Cursor or Codex # nodus mcp install > Add the Nodus MCP server to Claude Code, Cursor or Codex ## nodus mcp install [Section titled “nodus mcp install”](#nodus-mcp-install) Add the Nodus MCP server to Claude Code, Cursor or Codex ### Synopsis [Section titled “Synopsis”](#synopsis) Writes the nodus server into the client’s user-level configuration file and leaves everything else in it as it was. CLIENT is claude, cursor or codex. By default the client launches `nodus mcp` (local, acts with this machine’s login); with –hosted it connects to the hosted server and signs in with your browser instead. The original file is kept beside it as FILE.nodus-backup. A different existing nodus entry is not replaced unless you pass –force. ```plaintext nodus mcp install CLIENT [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus mcp install claude nodus mcp install cursor --hosted nodus mcp install codex --dry-run ``` ### Options [Section titled “Options”](#options) ```plaintext --dry-run Show what would change without writing --force Replace a different existing nodus entry -h, --help help for install --hosted Connect to the hosted server instead of launching nodus mcp --url string Hosted server URL (default: this context's server plus /mcp) ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus mcp](/docs/reference/cli/nodus_mcp/) - Serve Nodus tools to a local MCP client over stdio # nodus open > Open a resource's console page, or a port of its container, in the browser ## nodus open [Section titled “nodus open”](#nodus-open) Open a resource’s console page, or a port of its container, in the browser ### Synopsis [Section titled “Synopsis”](#synopsis) Without a port flag, open shows the resource’s console page. With –vscode, –jupyter or –port, it mints a preview of that container port (a single-use link valid for 60 seconds that signs the browser in for an hour) and opens it. ```plaintext nodus open KIND/NAME [--vscode | --jupyter | --port N] [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus open job/train nodus open ws/lab --vscode nodus open sb/dev --port 3000 --print ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for open --jupyter Open Jupyter (port 8888) --port int Open this container port --print Print the link instead of opening a browser --vscode Open VS Code in the browser (port 8080) ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus open job/train --print nodus open job/train nodus open sb/dev --vscode --print nodus open sb/dev --jupyter --print nodus open sb/dev --port 3000 nodus open job/train --port 6006 ``` # nodus patch > Update fields of a resource with a merge or JSON patch ## nodus patch [Section titled “nodus patch”](#nodus-patch) Update fields of a resource with a merge or JSON patch ```plaintext nodus patch KIND/NAME --patch PATCH [--type merge|json] [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus patch job/x --patch '{"spec":{"maxCostUSD":"50.00"}}' ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for patch --patch string The patch document (JSON) --type string merge or json (default "merge") ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus patch job/train --patch '{"spec":{"maxCostUSD":"50.00"}}' ``` # nodus pool > Pool utilization, forecasts, recommendations and actions ## nodus pool [Section titled “nodus pool”](#nodus-pool) Pool utilization, forecasts, recommendations and actions ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for pool ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus * [nodus pool approve](/docs/reference/cli/nodus_pool_approve/) - Approve a proposed PoolAction * [nodus pool forecast](/docs/reference/cli/nodus_pool_forecast/) - Show a pool’s demand forecast with p50 and p90 bands * [nodus pool pause](/docs/reference/cli/nodus_pool_pause/) - Pause every automated action in a pool * [nodus pool recommendations](/docs/reference/cli/nodus_pool_recommendations/) - List a pool’s recommendations * [nodus pool reject](/docs/reference/cli/nodus_pool_reject/) - Reject a proposed PoolAction * [nodus pool resume](/docs/reference/cli/nodus_pool_resume/) - Resume every automated action in a pool * [nodus pool revert](/docs/reference/cli/nodus_pool_revert/) - Undo an executed PoolAction where the action allows it * [nodus pool utilization](/docs/reference/cli/nodus_pool_utilization/) - Show a pool’s GPU utilization and device-hour ledger # nodus pool approve > Approve a proposed PoolAction ## nodus pool approve [Section titled “nodus pool approve”](#nodus-pool-approve) Approve a proposed PoolAction ```plaintext nodus pool approve ACTION [flags] ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for approve ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus pool](/docs/reference/cli/nodus_pool/) - Pool utilization, forecasts, recommendations and actions ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus pool approve lab-idle-1 ``` # nodus pool forecast > Show a pool's demand forecast with p50 and p90 bands ## nodus pool forecast [Section titled “nodus pool forecast”](#nodus-pool-forecast) Show a pool’s demand forecast with p50 and p90 bands ```plaintext nodus pool forecast POOL [flags] ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for forecast --horizon string Forecast horizon: 24h or 7d (default 7d) --json Print the response as JSON ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus pool](/docs/reference/cli/nodus_pool/) - Pool utilization, forecasts, recommendations and actions ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus pool forecast lab --horizon 7d nodus pool forecast missing ``` # nodus pool pause > Pause every automated action in a pool ## nodus pool pause [Section titled “nodus pool pause”](#nodus-pool-pause) Pause every automated action in a pool ```plaintext nodus pool pause POOL [flags] ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for pause ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus pool](/docs/reference/cli/nodus_pool/) - Pool utilization, forecasts, recommendations and actions ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus pool pause lab ``` # nodus pool recommendations > List a pool's recommendations ## nodus pool recommendations [Section titled “nodus pool recommendations”](#nodus-pool-recommendations) List a pool’s recommendations ```plaintext nodus pool recommendations POOL [flags] ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for recommendations --json Print the response as JSON --state string Only Open, Dismissed, Actioned or Expired recommendations ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus pool](/docs/reference/cli/nodus_pool/) - Pool utilization, forecasts, recommendations and actions ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus pool recommendations lab nodus pool recommendations lab --json ``` # nodus pool reject > Reject a proposed PoolAction ## nodus pool reject [Section titled “nodus pool reject”](#nodus-pool-reject) Reject a proposed PoolAction ```plaintext nodus pool reject ACTION [flags] ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for reject ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus pool](/docs/reference/cli/nodus_pool/) - Pool utilization, forecasts, recommendations and actions # nodus pool resume > Resume every automated action in a pool ## nodus pool resume [Section titled “nodus pool resume”](#nodus-pool-resume) Resume every automated action in a pool ```plaintext nodus pool resume POOL [flags] ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for resume ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus pool](/docs/reference/cli/nodus_pool/) - Pool utilization, forecasts, recommendations and actions ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus pool resume lab ``` # nodus pool revert > Undo an executed PoolAction where the action allows it ## nodus pool revert [Section titled “nodus pool revert”](#nodus-pool-revert) Undo an executed PoolAction where the action allows it ```plaintext nodus pool revert ACTION [flags] ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for revert ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus pool](/docs/reference/cli/nodus_pool/) - Pool utilization, forecasts, recommendations and actions ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus pool revert pact/lab-idle-1 ``` # nodus pool utilization > Show a pool's GPU utilization and device-hour ledger ## nodus pool utilization [Section titled “nodus pool utilization”](#nodus-pool-utilization) Show a pool’s GPU utilization and device-hour ledger ```plaintext nodus pool utilization POOL [flags] ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for utilization --json Print the response as JSON --window string Window to summarize, up to 30d (default 24h) ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus pool](/docs/reference/cli/nodus_pool/) - Pool utilization, forecasts, recommendations and actions ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus pool utilization lab ``` # nodus port-forward > Forward local ports to a Job, Sandbox or Workspace ## nodus port-forward [Section titled “nodus port-forward”](#nodus-port-forward) Forward local ports to a Job, Sandbox or Workspace ```plaintext nodus port-forward KIND/NAME LOCAL[:REMOTE]... [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus port-forward job/train 6006 nodus port-forward sb/dev 8080:80 ``` ### Options [Section titled “Options”](#options) ```plaintext --address string Local address to listen on (default "127.0.0.1") -h, --help help for port-forward --index int Job index (Indexed Jobs) (default -1) --rank string Gang member rank (default 0) ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus # nodus request > Request an action: restart, refresh-secrets, reload-sinks, verify, resend, start, reimport, approve-burst, drain ## nodus request [Section titled “nodus request”](#nodus-request) Request an action: restart, refresh-secrets, reload-sinks, verify, resend, start, reimport, approve-burst, drain ```plaintext nodus request ACTION KIND/NAME [flags] ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for request ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus request restart job/train ``` # nodus resume > Resume a suspended run ## nodus resume [Section titled “nodus resume”](#nodus-resume) Resume a suspended run ```plaintext nodus resume KIND/NAME [flags] ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for resume ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus resume job/train ``` # nodus rollout > Roll out changes ## nodus rollout [Section titled “nodus rollout”](#nodus-rollout) Roll out changes ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for rollout ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus * [nodus rollout restart](/docs/reference/cli/nodus_rollout_restart/) - Restart without a spec change (alias of `request restart`) # nodus rollout restart > Restart without a spec change (alias of `request restart`) ## nodus rollout restart [Section titled “nodus rollout restart”](#nodus-rollout-restart) Restart without a spec change (alias of `request restart`) ```plaintext nodus rollout restart KIND/NAME [flags] ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for restart ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus rollout](/docs/reference/cli/nodus_rollout/) - Roll out changes ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus rollout restart job/train ``` # nodus run > Run a command as a Job: estimate, phases, logs and the final cost ## nodus run [Section titled “nodus run”](#nodus-run) Run a command as a Job: estimate, phases, logs and the final cost ### Synopsis [Section titled “Synopsis”](#synopsis) Uploads the current directory (unless –no-source), creates a Job, prints its estimate and phases, streams its logs and ends with the cost line. The command’s exit code becomes nodus’s; 125 means a Nodus or API error, 124 a timeout and 130 an interrupt. On a terminal Ctrl+C asks before canceling; -d detaches. ```plaintext nodus run [flags] -- COMMAND [ARGS...] | run APP.py[::FUNCTION] [ARGS...] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus run --gpu L4 --image nodus/pytorch -- python train.py nodus run --gpu H100:8 --max-cost 40 -d -- torchrun train.py nodus run app.py::main ``` ### Options [Section titled “Options”](#options) ```plaintext --allow-large-source Upload source above 500 MiB compressed --checkpoint string Checkpoint path (default /nodus/state) --complete-by string Finish-by time (RFC 3339) for placement --completions int Indexed completions --continuity string Recovery: Checkpointed, Restartable or Ephemeral --cpu string vCPUs -d, --detach Return once the Job is created --disk string Disk, such as 200Gi --dry-run Print the estimate only --env stringArray Environment variable K=V --expected-duration duration Expected runtime, for the estimate --gpu string Accelerator TYPE[:COUNT], such as L4 or H100:8 --gpus-per-node int Beta: GPUs per gang node -h, --help help for run --idempotency-key string Override the generated idempotency key --image string Image, such as nodus/pytorch --interruptible string[="prefer"] Interruptible capacity: allow, prefer or never --keep Keep the Job after it finishes (no 30-day TTL) -l, --label strings Labels k=v --launcher string Beta: plain, torchrun, ray or verl --max-cost string Spend cap in USD --memory string Memory, such as 64Gi --name string Job name (default: generated run-xxxxx) --network string Beta: colocated, regional or global --no-source Do not upload the current directory --no-warm Release capacity at once instead of keeping it warm for 60 s --nodes int Beta: gang nodes --output stringArray Output NAME=/path --parallelism int Indexes running at once --profile string Placement profile: Balanced, Cost or Speed --region string Region class, such as us or eu --secret stringArray Secret to expose as environment variables --startup-timeout duration Beta: gang assembly budget --timeout duration Wall-clock limit, such as 2h --total-gpus int Beta: total GPUs across the gang --transport string Beta: direct or auto --volume stringArray Volume NAME:/path[:ro] ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus run app.py::main --epochs 3 nodus run fail.py nodus run --gpu L4 --image nodus/pytorch -- python train.py nodus run --gpu L4 --no-source -- sh -c 'exit 3' nodus run --gpu L4 --no-source --name late -- python deadline.py nodus run -d --no-source --gpu L4 --keep --name detached -- python train.py nodus run --dry-run --no-source --gpu L4 -- python train.py nodus run -d --no-source --name gang --gpu H100 --gpus-per-node 8 --nodes 2 --launcher torchrun --network regional --max-cost 40 --env A=1 --interruptible -- torchrun train.py nodus run --dry-run --no-source --gpu H100 --gpus-per-node 8 --nodes 2 -- torchrun train.py nodus run --no-source --gpu none -- python train.py ``` # nodus secret > Create Secrets from literals, env files or registry logins ## nodus secret [Section titled “nodus secret”](#nodus-secret) Create Secrets from literals, env files or registry logins ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for secret ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus * [nodus secret create](/docs/reference/cli/nodus_secret_create/) - Create a Secret # nodus secret create > Create a Secret ## nodus secret create [Section titled “nodus secret create”](#nodus-secret-create) Create a Secret ```plaintext nodus secret create NAME [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus secret create hf-token --from-literal HF_TOKEN=hf_xxx nodus secret create app-env --from-env-file .env nodus secret create registry --type Registry --from-literal server=ghcr.io --from-literal username=ada --from-literal password=… ``` ### Options [Section titled “Options”](#options) ```plaintext --dry-run string[="server"] server returns the estimate without creating; client prints the manifest (default "none") --from-env-file stringArray File of KEY=VALUE lines --from-literal stringArray KEY=VALUE (repeatable) -h, --help help for create -o, --output string Output format: name, json, yaml, or estimate (with --dry-run=server) --type string Opaque or Registry ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus secret](/docs/reference/cli/nodus_secret/) - Create Secrets from literals, env files or registry logins # nodus serve > Run the App in APP.py and redeploy it whenever a file changes ## nodus serve [Section titled “nodus serve”](#nodus-serve) Run the App in APP.py and redeploy it whenever a file changes ```plaintext nodus serve APP.py [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus serve app.py ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for serve ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus serve app.py ``` # nodus shell > Open an interactive shell in a Sandbox, creating it if it is not there ## nodus shell [Section titled “nodus shell”](#nodus-shell) Open an interactive shell in a Sandbox, creating it if it is not there ### Synopsis [Section titled “Synopsis”](#synopsis) shell opens a terminal in a Sandbox. A name that does not exist yet is created from the same flags as `nodus create sandbox` (image nodus/agent-tools unless –image says otherwise), and the Sandbox stays after the shell exits, so the next `nodus shell` finds /workspace as you left it; –rm deletes a Sandbox this command created when the shell exits. Without a name the shell gets a Sandbox of its own, which is deleted when it exits unless –keep is set. The command after – replaces bash. ```plaintext nodus shell [sandbox/NAME] [flags] [-- COMMAND...] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus shell sandbox/dev nodus shell sandbox/dev --cpu 4 --memory 8Gi nodus shell sandbox/dev --rm nodus shell nodus shell sandbox/dev -- zsh ``` ### Options [Section titled “Options”](#options) ```plaintext --cpu string vCPU floor --disk string Disk floor, such as 50Gi --gpu string Accelerator TYPE[:COUNT], such as L4 or H100:8 -h, --help help for shell --idle-timeout string Stop after this idle time --image string Container image (default "nodus/agent-tools") --keep Without a name: keep the Sandbox after the shell exits --max-cost string Spend limit in USD --memory string Memory floor, such as 16Gi --rm Delete the Sandbox when the shell exits (only one this command created) --secret stringArray Secret to mount (repeatable) ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus shell sandbox/lab nodus shell sandbox/lab -- zsh nodus shell sandbox/scratch --rm nodus shell ``` # nodus ssh > Connect to a Workspace over SSH ## nodus ssh [Section titled “nodus ssh”](#nodus-ssh) Connect to a Workspace over SSH ### Synopsis [Section titled “Synopsis”](#synopsis) ssh uses the account-bound connection rule installed by nodus login for ...nodus (so plain ssh, scp, rsync and VS Code Remote-SSH work too) and connects. A Workspace stopped by idle, schedule or credits starts when you connect; one you stopped needs `nodus start`. Sign in once per computer with nodus login to create a local SSH key and register only its public key. Use –setup to repair SSH enrollment without starting or connecting to any Workspace. ```plaintext nodus ssh workspace/NAME [-- COMMAND...] [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus login nodus ssh workspace/lab nodus ssh workspace/lab -- nvidia-smi ``` ### Options [Section titled “Options”](#options) ```plaintext --config Print the ~/.ssh/config entry instead of connecting -h, --help help for ssh --setup Repair this account's SSH setup for all Workspaces without starting or connecting ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus # nodus start > Start a stopped Sandbox, Workspace, Function, Agent or InferenceEndpoint ## nodus start [Section titled “nodus start”](#nodus-start) Start a stopped Sandbox, Workspace, Function, Agent or InferenceEndpoint ```plaintext nodus start KIND/NAME [flags] ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for start ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus start sb/dev ``` # nodus stop > Stop a Sandbox, Workspace, Function, Agent or InferenceEndpoint ## nodus stop [Section titled “nodus stop”](#nodus-stop) Stop a Sandbox, Workspace, Function, Agent or InferenceEndpoint ```plaintext nodus stop KIND/NAME [flags] ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for stop ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus stop sb/dev nodus stop sb/missing ``` # nodus suspend > Suspend a Job, Pipeline, Sweep, TrainingJob or AgentRun (checkpointing first) ## nodus suspend [Section titled “nodus suspend”](#nodus-suspend) Suspend a Job, Pipeline, Sweep, TrainingJob or AgentRun (checkpointing first) ```plaintext nodus suspend KIND/NAME [flags] ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for suspend ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus suspend job/train ``` # nodus top > Show GPU, CPU and memory use, the hourly rate and spend of running Jobs ## nodus top [Section titled “nodus top”](#nodus-top) Show GPU, CPU and memory use, the hourly rate and spend of running Jobs ```plaintext nodus top jobs|sandboxes|workspaces|functions [NAME] [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus top jobs nodus top sandbox dev ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for top --no-headers Omit table headers ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus top jobs nodus top job idle nodus top jobs --no-headers nodus top sandboxes nodus top job missing ``` # nodus version > Print the client version ## nodus version [Section titled “nodus version”](#nodus-version) Print the client version ```plaintext nodus version [flags] ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for version ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus version ``` # nodus volume > Put, get, list and remove the files of a Volume ## nodus volume [Section titled “nodus volume”](#nodus-volume) Put, get, list and remove the files of a Volume ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for volume ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus * [nodus volume clear](/docs/reference/cli/nodus_volume_clear/) - Commit an empty revision; the Volume and its history stay * [nodus volume get](/docs/reference/cli/nodus_volume_get/) - Download a file or directory from a Volume * [nodus volume ls](/docs/reference/cli/nodus_volume_ls/) - List a directory of a Volume * [nodus volume put](/docs/reference/cli/nodus_volume_put/) - Add a local file or directory to a Volume as a new revision * [nodus volume rm](/docs/reference/cli/nodus_volume_rm/) - Remove a file or directory from a Volume as a new revision # nodus volume clear > Commit an empty revision; the Volume and its history stay ## nodus volume clear [Section titled “nodus volume clear”](#nodus-volume-clear) Commit an empty revision; the Volume and its history stay ```plaintext nodus volume clear NAME [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus volume clear home ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for clear ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus volume](/docs/reference/cli/nodus_volume/) - Put, get, list and remove the files of a Volume # nodus volume get > Download a file or directory from a Volume ## nodus volume get [Section titled “nodus volume get”](#nodus-volume-get) Download a file or directory from a Volume ### Synopsis [Section titled “Synopsis”](#synopsis) Downloads REMOTE from the latest revision (or –revision) into LOCAL (default .). LOCAL - writes a file to standard output. Existing local files are kept unless –force. ```plaintext nodus volume get NAME REMOTE [LOCAL] [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus volume get data /train/part-0.parquet ./ nodus volume get data /train ./train --revision 3 nodus volume get data /config.yaml - ``` ### Options [Section titled “Options”](#options) ```plaintext --force Overwrite existing local files -h, --help help for get --revision int32 Revision to read (default the latest) ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus volume](/docs/reference/cli/nodus_volume/) - Put, get, list and remove the files of a Volume # nodus volume ls > List a directory of a Volume ## nodus volume ls [Section titled “nodus volume ls”](#nodus-volume-ls) List a directory of a Volume ```plaintext nodus volume ls NAME [PATH] [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus volume ls data /train nodus volume ls data --revision 2 ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for ls --revision int32 Revision to list (default the latest) ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus volume](/docs/reference/cli/nodus_volume/) - Put, get, list and remove the files of a Volume # nodus volume put > Add a local file or directory to a Volume as a new revision ## nodus volume put [Section titled “nodus volume put”](#nodus-volume-put) Add a local file or directory to a Volume as a new revision ### Synopsis [Section titled “Synopsis”](#synopsis) Adds LOCAL at REMOTE (default /) on top of the latest revision; everything else in the Volume stays. A directory’s contents merge into REMOTE; a file lands at REMOTE, or inside it when REMOTE is a directory or ends in /. –extract expands a .zip, .tar, .tar.gz, .tgz or .tar.zst archive into REMOTE instead. ```plaintext nodus volume put NAME LOCAL [REMOTE] [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus volume put data ./dataset /train nodus volume put data ./config.yaml /etc/ nodus volume put data ./archive.tar.gz /raw --extract ``` ### Options [Section titled “Options”](#options) ```plaintext --extract Expand the LOCAL archive into REMOTE -h, --help help for put ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus volume](/docs/reference/cli/nodus_volume/) - Put, get, list and remove the files of a Volume # nodus volume rm > Remove a file or directory from a Volume as a new revision ## nodus volume rm [Section titled “nodus volume rm”](#nodus-volume-rm) Remove a file or directory from a Volume as a new revision ```plaintext nodus volume rm NAME PATH [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus volume rm data /train/stale.csv ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for rm ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus volume](/docs/reference/cli/nodus_volume/) - Put, get, list and remove the files of a Volume # nodus wait > Wait until a resource meets a condition ## nodus wait [Section titled “nodus wait”](#nodus-wait) Wait until a resource meets a condition ```plaintext nodus wait KIND/NAME --for=condition=C|jsonpath='{...}'=VALUE|delete [flags] ``` ### Examples [Section titled “Examples”](#examples) ```plaintext nodus wait job/x --for=jsonpath='{.status.phase}'=Succeeded --timeout 1h ``` ### Options [Section titled “Options”](#options) ```plaintext --for string condition=NAME[=STATUS], jsonpath='{.path}'=VALUE, or delete -h, --help help for wait --timeout duration Give up after this long (exit 124) (default 30m0s) ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus wait job/train --for=jsonpath={.status.phase}=Succeeded --timeout 1m nodus wait job/train --for=condition=Scheduled nodus wait job/train --for=jsonpath={.status.phase}=Running ``` # nodus whoami > Show the signed-in principal, org, role, project, scopes and available credit ## nodus whoami [Section titled “nodus whoami”](#nodus-whoami) Show the signed-in principal, org, role, project, scopes and available credit ```plaintext nodus whoami [flags] ``` ### Options [Section titled “Options”](#options) ```plaintext -h, --help help for whoami -o, --output string json for machine-readable output ``` ### Options inherited from parent commands [Section titled “Options inherited from parent commands”](#options-inherited-from-parent-commands) ```plaintext --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials) ``` ### SEE ALSO [Section titled “SEE ALSO”](#see-also) * [nodus](/docs/reference/cli/nodus/) - Run compute Jobs on Nodus ### Tested examples [Section titled “Tested examples”](#tested-examples) These invocations run in CI against the CLI’s fake API. ```sh nodus whoami ``` # Environments > Every catalog Environment, its category, reward and readiness. The catalog Environments in project `nodus`; reference one as `spec.environment.name: nodus/@` on a TrainingJob or an agent eval. | Environment | Category | Reward | Readiness | Summary | | ------------------------------------------------------------------------------------------ | --------- | ------ | --------- | ---------------------------------------------------------------------------------------------------------- | | [`arithmetic-v2`](/docs/reference/environments/arithmetic-v2/) | Math | Binary | Stable | Evaluate short integer expressions with standard precedence and give the exact result. | | [`graph-coloring`](/docs/reference/environments/graph-coloring/) | Reasoning | Binary | Stable | Color a small graph with three colors so that no edge joins two nodes of the same color. | | [`gsm8k`](/docs/reference/environments/gsm8k/) | Math | Binary | Research | Eight thousand grade-school maths word problems, each with a worked solution and a single numeric answer. | | [`letter-counting-legacy-eval`](/docs/reference/environments/letter-counting-legacy-eval/) | Reasoning | Scalar | Research | Legacy letter-counting evaluation protocol: first 64 generated tasks, raw completion scoring. | | [`letter-counting-legacy-rl`](/docs/reference/environments/letter-counting-legacy-rl/) | Reasoning | Scalar | Research | Legacy letter-counting RL protocol: first 64 generated tasks train, next 64 held out. | | [`python-functions`](/docs/reference/environments/python-functions/) | Code | Binary | Stable | Write solve(values: list\[int]) -> int from a one-line specification; private cases decide the verdict. | | [`reasoning-gym`](/docs/reference/environments/reasoning-gym/) | Reasoning | Scalar | Research | Procedural reasoning generators with deterministic algorithmic scorers; only reviewed families are served. | # arithmetic-v2 > Evaluate short integer expressions with standard precedence and give the exact result. Evaluate short integer expressions with standard precedence and give the exact result. | Field | Value | | ----------------- | ------------------------------------------------------------------------------- | | Reference | `nodus/arithmetic-v2@2.0.0` | | Image | `nodus/env-arithmetic-v2:2.0.0` | | Publisher | Nodus | | Category | Math | | Readiness | Stable | | Modes | Train, Evaluate | | Reward | Binary | | Held-out measures | TrainedTask | | Splits | 500 train, 128 test (disjoint by canonical identity) | | Licenses | code Apache-2.0, data Apache-2.0 | | Source | | ## Grading [Section titled “Grading”](#grading) Completions are graded by the platform, never by the trainer: the trainer submits `{taskId, completion}` batches and the verdicts come back as task events. | Grader | Kind | What it checks | | --------------- | ------- | ---------------------------------------- | | `exact-integer` | Program | The last N equals the expression’s value | ## Sample task [Section titled “Sample task”](#sample-task) One line of `nodus-env tasks --split test --seed 0`; tasks never carry answers. ```json { "metadata": { "terms": 3 }, "prompt": "Compute 41 - 38 * 33. Give only the final integer, like \u003canswer\u003e42\u003c/answer\u003e.", "taskId": "arithmetic-v2:test:0:0" } ``` ## Examples [Section titled “Examples”](#examples) Each example is a TrainingJob template. A baseline is shown only where the example was measured by running it. | Example | Mode | Runtime | Model | Tasks | Baseline | Trained | Measured on | | ----------------- | ----- | ----------------- | --------------------------- | ----- | ------------ | ------------ | ------------ | | `arithmetic-grpo` | Train | `nodus/grpo-lora` | Qwen/Qwen3-0.6B @ `c1899de` | 64 | not measured | not measured | not measured | Run one with a server dry-run first: ```console $ nodus create trainingjob my-run --from-example nodus/arithmetic-v2:arithmetic-grpo --dry-run=server -o estimate ``` ```python import nodus job = nodus.recipes.TrainingJob.from_example("nodus/arithmetic-v2:arithmetic-grpo") plan = job.preview() run = plan.run(max_cost=5) print(run.wait().summary) ``` # graph-coloring > Color a small graph with three colors so that no edge joins two nodes of the same color. Color a small graph with three colors so that no edge joins two nodes of the same color. | Field | Value | | ----------------- | ------------------------------------------------------------------------------- | | Reference | `nodus/graph-coloring@1.0.0` | | Image | `nodus/env-graph-coloring:1.0.0` | | Publisher | Nodus | | Category | Reasoning | | Readiness | Stable | | Modes | Train, Evaluate | | Reward | Binary | | Held-out measures | TrainedTask | | Splits | 200 train, 64 test (disjoint by canonical identity) | | Licenses | code Apache-2.0, data Apache-2.0 | | Source | | ## Grading [Section titled “Grading”](#grading) Completions are graded by the platform, never by the trainer: the trainer submits `{taskId, completion}` batches and the verdicts come back as task events. | Grader | Kind | What it checks | | ----------------- | ------- | ---------------------------------------------------------------------------- | | `proper-coloring` | Program | Checks every edge joins two different colors; any proper coloring is correct | ## Sample task [Section titled “Sample task”](#sample-task) One line of `nodus-env tasks --split test --seed 0`; tasks never carry answers. ```json { "metadata": { "edges": 12, "nodes": 7 }, "prompt": "A graph has 7 nodes numbered 0 to 6 and these edges: 0-1, 0-2, 0-4, 1-3, 1-6, 2-3, 2-4, 2-5, 3-4, 3-5, 3-6, 4-6.\nAssign each node one of the colors 0, 1 or 2 so that no edge connects two nodes of the same color.\nAnswer with the colors of nodes 0 to 6 in order, like \u003canswer\u003e[0, 1, 2, 0]\u003c/answer\u003e.", "taskId": "graph-coloring:test:0:0" } ``` ## Examples [Section titled “Examples”](#examples) Each example is a TrainingJob template. A baseline is shown only where the example was measured by running it. | Example | Mode | Runtime | Model | Tasks | Baseline | Trained | Measured on | | --------------------- | ----- | ----------------- | --------------------------- | ----- | ------------ | ------------ | ------------ | | `graph-coloring-grpo` | Train | `nodus/grpo-lora` | Qwen/Qwen3-0.6B @ `c1899de` | 64 | not measured | not measured | not measured | Run one with a server dry-run first: ```console $ nodus create trainingjob my-run --from-example nodus/graph-coloring:graph-coloring-grpo --dry-run=server -o estimate ``` ```python import nodus job = nodus.recipes.TrainingJob.from_example("nodus/graph-coloring:graph-coloring-grpo") plan = job.preview() run = plan.run(max_cost=5) print(run.wait().summary) ``` # gsm8k > Eight thousand grade-school maths word problems, each with a worked solution and a single numeric answer. Eight thousand grade-school maths word problems, each with a worked solution and a single numeric answer. | Field | Value | | ----------------- | -------------------------------------------------------------------------------------------- | | Reference | `nodus/gsm8k@1.0.0` | | Image | `nodus/env-gsm8k:1.0.0` | | Publisher | OpenAI | | Category | Math | | Readiness | Research | | Modes | Train, Evaluate | | Reward | Binary | | Held-out measures | TrainedTask | | Splits | 7473 train, 1319 test (disjoint by canonical identity) | | Licenses | code MIT, data MIT | | Source | | ## Grading [Section titled “Grading”](#grading) Completions are graded by the platform, never by the trainer: the trainer submits `{taskId, completion}` batches and the verdicts come back as task events. | Grader | Kind | What it checks | | -------------- | ------- | ------------------------------------------------------------------ | | `final-number` | Program | The last number, else the last number, equals the value after #### | ## Sample task [Section titled “Sample task”](#sample-task) One line of `nodus-env tasks --split test --seed 0`; tasks never carry answers. ```json { "metadata": { "source": "openai/gsm8k" }, "prompt": "A fruit vendor bought 50 watermelons for $80. He sold all of them at a profit of 25%. How much was each watermelon sold?\n\nWork through the problem step by step, then give the final number inside answer tags, like \u003canswer\u003e42\u003c/answer\u003e.", "taskId": "gsm8k:test:0:0" } ``` ## Examples [Section titled “Examples”](#examples) Each example is a TrainingJob template. A baseline is shown only where the example was measured by running it. | Example | Mode | Runtime | Model | Tasks | Baseline | Trained | Measured on | | --------------- | ----- | ----------------- | --------------------------- | ----- | -------- | ------- | ---------------------- | | `gsm8k-trained` | Train | `nodus/grpo-lora` | Qwen/Qwen3-1.7B @ `70d244c` | 64 | 79.7 % | 87.5 % | a40-48g-x1, 2026-09-24 | Run one with a server dry-run first: ```console $ nodus create trainingjob my-run --from-example nodus/gsm8k:gsm8k-trained --dry-run=server -o estimate ``` ```python import nodus job = nodus.recipes.TrainingJob.from_example("nodus/gsm8k:gsm8k-trained") plan = job.preview() run = plan.run(max_cost=5) print(run.wait().summary) ``` # letter-counting-legacy-eval > Legacy letter-counting evaluation protocol: first 64 generated tasks, raw completion scoring. Legacy letter-counting evaluation protocol: first 64 generated tasks, raw completion scoring. | Field | Value | | ----------------- | ------------------------------------------------- | | Reference | `nodus/letter-counting-legacy-eval@1.0.0` | | Image | `nodus/env-letter-counting-legacy-eval:1.0.0` | | Publisher | Open-Thought | | Category | Reasoning | | Readiness | Research | | Modes | Evaluate | | Reward | Scalar | | Held-out measures | TrainedTask | | Splits | 0 train, 64 test (disjoint by canonical identity) | | Licenses | code Apache-2.0, data Apache-2.0 | | Source | | ## Grading [Section titled “Grading”](#grading) Completions are graded by the platform, never by the trainer: the trainer submits `{taskId, completion}` batches and the verdicts come back as task events. | Grader | Kind | What it checks | | -------------- | ------- | -------------------------------------------------------------------------- | | `score-answer` | Program | Raw completion scored by pinned Reasoning Gym; only full credit is Correct | ## Sample task [Section titled “Sample task”](#sample-task) One line of `nodus-env tasks --split test --seed 0`; tasks never carry answers. ```json { "metadata": { "family": "letter_counting", "systemPrompt": "Reply with the final answer only. Do not restate the question, do not show working, and do not add any words, labels or punctuation around it." }, "prompt": "How many times does the letter \"b\" appear in the text: \"phrase Project Gutenberg associated with or appearing on the work you\"?", "taskId": "letter-counting-legacy-eval:test:0:0" } ``` ## Examples [Section titled “Examples”](#examples) Each example is a TrainingJob template. A baseline is shown only where the example was measured by running it. | Example | Mode | Runtime | Model | Tasks | Baseline | Trained | Measured on | | ----------------- | -------- | ---------------- | --------------------------- | ----- | ------------ | ---------------- | ------------ | | `letter-counting` | Evaluate | `nodus/evaluate` | Qwen/Qwen3-1.7B @ `70d244c` | 64 | not measured | n/a (evaluation) | not measured | Run one with a server dry-run first: ```console $ nodus create trainingjob my-run --from-example nodus/letter-counting-legacy-eval:letter-counting --dry-run=server -o estimate ``` ```python import nodus job = nodus.recipes.TrainingJob.from_example("nodus/letter-counting-legacy-eval:letter-counting") plan = job.preview() run = plan.run(max_cost=5) print(run.wait().summary) ``` # letter-counting-legacy-rl > Legacy letter-counting RL protocol: first 64 generated tasks train, next 64 held out. Legacy letter-counting RL protocol: first 64 generated tasks train, next 64 held out. | Field | Value | | ----------------- | -------------------------------------------------- | | Reference | `nodus/letter-counting-legacy-rl@1.0.0` | | Image | `nodus/env-letter-counting-legacy-rl:1.0.0` | | Publisher | Open-Thought | | Category | Reasoning | | Readiness | Research | | Modes | Train, Evaluate | | Reward | Scalar | | Held-out measures | TrainedTask | | Splits | 64 train, 64 test (disjoint by canonical identity) | | Licenses | code Apache-2.0, data Apache-2.0 | | Source | | ## Grading [Section titled “Grading”](#grading) Completions are graded by the platform, never by the trainer: the trainer submits `{taskId, completion}` batches and the verdicts come back as task events. | Grader | Kind | What it checks | | -------------- | ------- | -------------------------------------------------------------------------- | | `score-answer` | Program | Raw completion scored by pinned Reasoning Gym; only full credit is Correct | ## Sample task [Section titled “Sample task”](#sample-task) One line of `nodus-env tasks --split test --seed 0`; tasks never carry answers. ```json { "metadata": { "family": "letter_counting", "systemPrompt": "Reply with the final answer only. Do not restate the question, do not show working, and do not add any words, labels or punctuation around it." }, "prompt": "How many times does the letter \"t\" appear in the text: \"are far more complex than all that In real life every act\"?", "taskId": "letter-counting-legacy-rl:test:0:0" } ``` ## Examples [Section titled “Examples”](#examples) Each example is a TrainingJob template. A baseline is shown only where the example was measured by running it. | Example | Mode | Runtime | Model | Tasks | Baseline | Trained | Measured on | | ----------------- | ----- | ----------------- | --------------------------- | ----- | ------------ | ------------ | ------------ | | `letter-counting` | Train | `nodus/grpo-lora` | Qwen/Qwen3-1.7B @ `70d244c` | 64 | not measured | not measured | not measured | Run one with a server dry-run first: ```console $ nodus create trainingjob my-run --from-example nodus/letter-counting-legacy-rl:letter-counting --dry-run=server -o estimate ``` ```python import nodus job = nodus.recipes.TrainingJob.from_example("nodus/letter-counting-legacy-rl:letter-counting") plan = job.preview() run = plan.run(max_cost=5) print(run.wait().summary) ``` # python-functions > Write solve(values: list[int]) -> int from a one-line specification; private cases decide the verdict. Write solve(values: list\[int]) -> int from a one-line specification; private cases decide the verdict. | Field | Value | | ----------------- | ------------------------------------------------------------------------------- | | Reference | `nodus/python-functions@1.0.0` | | Image | `nodus/env-python-functions:1.0.0` | | Publisher | Nodus | | Category | Code | | Readiness | Stable | | Modes | Train, Evaluate | | Reward | Binary | | Held-out measures | TrainedTask | | Splits | 20 train, 16 test (disjoint by canonical identity) | | Licenses | code Apache-2.0, data Apache-2.0 | | Source | | ## Grading [Section titled “Grading”](#grading) Completions are graded by the platform, never by the trainer: the trainer submits `{taskId, completion}` batches and the verdicts come back as task events. | Grader | Kind | What it checks | | --------------- | ------- | ---------------------------------------------------------------------------------------------------- | | `private-cases` | Program | Runs the candidate as an unprivileged uid on public inputs; every private expected output must match | Grading runs in a sandbox (`nodus/env-python-functions:1.0.0`, 1 CPU, 1Gi memory, 30s timeout, no network); candidate code runs as an unprivileged user that cannot read the expected answers. ## Sample task [Section titled “Sample task”](#sample-task) One line of `nodus-env tasks --split test --seed 0`; tasks never carry answers. ```json { "metadata": { "function": "count-divisible-by-last" }, "prompt": "Write Python source defining solve(values: list[int]) -\u003e int. Return how many values are divisible by the last value, or 0 when it is zero. Return 0 for an empty list. Output only Python source, without Markdown fences.", "taskId": "python-functions:test:0:0" } ``` ## Examples [Section titled “Examples”](#examples) Each example is a TrainingJob template. A baseline is shown only where the example was measured by running it. | Example | Mode | Runtime | Model | Tasks | Baseline | Trained | Measured on | | ----------------------- | ----- | ----------------- | --------------------------- | ----- | ------------ | ------------ | ------------ | | `python-functions-grpo` | Train | `nodus/grpo-lora` | Qwen/Qwen3-0.6B @ `c1899de` | 16 | not measured | not measured | not measured | Run one with a server dry-run first: ```console $ nodus create trainingjob my-run --from-example nodus/python-functions:python-functions-grpo --dry-run=server -o estimate ``` ```python import nodus job = nodus.recipes.TrainingJob.from_example("nodus/python-functions:python-functions-grpo") plan = job.preview() run = plan.run(max_cost=5) print(run.wait().summary) ``` # reasoning-gym > Procedural reasoning generators with deterministic algorithmic scorers; only reviewed families are served. Procedural reasoning generators with deterministic algorithmic scorers; only reviewed families are served. | Field | Value | | ----------------- | ----------------------------------------------------- | | Reference | `nodus/reasoning-gym@1.0.0` | | Image | `nodus/env-reasoning-gym:1.0.0` | | Publisher | Open-Thought | | Category | Reasoning | | Readiness | Research | | Modes | Train, Evaluate | | Reward | Scalar | | Held-out measures | TrainedTask | | Splits | 2000 train, 320 test (disjoint by canonical identity) | | Licenses | code Apache-2.0, data Apache-2.0 | | Source | | ## Grading [Section titled “Grading”](#grading) Completions are graded by the platform, never by the trainer: the trainer submits `{taskId, completion}` batches and the verdicts come back as task events. | Grader | Kind | What it checks | | -------------- | ------- | -------------------------------------------------------------------------------- | | `score-answer` | Program | The family’s score_answer; only a full score is Correct, the score is the reward | ## Sample task [Section titled “Sample task”](#sample-task) One line of `nodus-env tasks --split test --seed 0`; tasks never carry answers. ```json { "metadata": { "family": "basic_arithmetic" }, "prompt": "Calculate -3 - 8 / ( -8 + 1 + 8 ).\n\nReply with the final answer only. Do not restate the question, do not show working, and do not add any words, labels or punctuation around it.", "taskId": "reasoning-gym:test:0:0" } ``` ## Examples [Section titled “Examples”](#examples) Each example is a TrainingJob template. A baseline is shown only where the example was measured by running it. | Example | Mode | Runtime | Model | Tasks | Baseline | Trained | Measured on | | ------------------ | -------- | ---------------- | --------------------------- | ----- | ------------ | ---------------- | ------------ | | `chain-sum` | Evaluate | `nodus/evaluate` | Qwen/Qwen3-1.7B @ `70d244c` | 64 | not measured | n/a (evaluation) | not measured | | `number-format` | Evaluate | `nodus/evaluate` | Qwen/Qwen3-1.7B @ `70d244c` | 64 | not measured | n/a (evaluation) | not measured | | `basic-arithmetic` | Evaluate | `nodus/evaluate` | Qwen/Qwen3-1.7B @ `70d244c` | 64 | not measured | n/a (evaluation) | not measured | | `products` | Evaluate | `nodus/evaluate` | Qwen/Qwen3-1.7B @ `70d244c` | 64 | not measured | n/a (evaluation) | not measured | | `letter-counting` | Evaluate | `nodus/evaluate` | Qwen/Qwen3-1.7B @ `70d244c` | 64 | not measured | n/a (evaluation) | not measured | Run one with a server dry-run first: ```console $ nodus create trainingjob my-run --from-example nodus/reasoning-gym:chain-sum --dry-run=server -o estimate ``` ```python import nodus job = nodus.recipes.TrainingJob.from_example("nodus/reasoning-gym:chain-sum") plan = job.preview() run = plan.run(max_cost=5) print(run.wait().summary) ``` # Error codes > One page per API error code, with its HTTP status, what it means and how to fix it. Every API error is a Kubernetes-style `Status` object. `reason` is a stable code, `code` is the HTTP status, and three extra fields tell you what to do next: ```json { "kind": "Status", "apiVersion": "v1", "status": "Failure", "reason": "InsufficientCredits", "code": 402, "message": "job \"train-a\" needs a $3.20 hold to start (released when it ends); available $1.10", "fix": "nodus billing top-up 20, or lower spec.maxCostUSD", "docs": "https://nodus-compute.ai/docs/reference/errors/insufficient-credits", "requestId": "req_01j9…" } ``` Include the `requestId` when you contact support. The `docs` link is the code’s page below, named after the code in lowercase words (`InsufficientCredits` is `insufficient-credits`). [AdmissionsPaused ](/docs/reference/errors/admissions-paused/)Operators paused admissions for this kind; existing objects keep running. [AgentGroupClosed ](/docs/reference/errors/agent-group-closed/)A sealed, canceled, finished or deleting AgentGroup admits no new member runs. [AgentGroupFull ](/docs/reference/errors/agent-group-full/)An AgentGroup admits at most spec.limits.maxPending unfinished member runs. [AgentRunFinished ](/docs/reference/errors/agent-run-finished/)A run in a terminal phase (Succeeded, Failed or canceled) no longer reads its inbox. [AgentStopped ](/docs/reference/errors/agent-stopped/)An Agent in the Stopped state refuses new runs; runs already created are not affected. [AlreadyExists ](/docs/reference/errors/already-exists/)A create by name found an object with that name whose spec differs; an identical create returns the object. [ArrearsOutstanding ](/docs/reference/errors/arrears-outstanding/)Charges that the org's credits could not cover are unpaid; they block new work until they are settled. [BadRequest ](/docs/reference/errors/bad-request/)The request is malformed: a query parameter, selector or body could not be parsed. [BudgetExceeded ](/docs/reference/errors/budget-exceeded/)A Block budget or a maxCostUSD cap that applies to the object has less room than the action's hold. [CapacityUnavailable ](/docs/reference/errors/capacity-unavailable/)No offering that meets the requirements has capacity at the moment. [CloudAccountNotVerified ](/docs/reference/errors/cloud-account-not-verified/)Nodus reads inventory only after it has verified the account's read-only access. [Conflict ](/docs/reference/errors/conflict/)The object changed after the caller read it (the resourceVersion is stale). [ConnectionNotReady ](/docs/reference/errors/connection-not-ready/)A Connection is usable once a verification passed. [EncryptionFailed ](/docs/reference/errors/encryption-failed/)The org's encryption key was unavailable; nothing was stored. [EnrollmentTokenInvalid ](/docs/reference/errors/enrollment-token-invalid/)An enrollment token works once, for 24 hours, for the pool it was created for. [Expired ](/docs/reference/errors/expired/)The watch position or list continue token is older than the 24-hour change history. [FieldImmutable ](/docs/reference/errors/field-immutable/)The update changes a spec field that is immutable after create; the causes list each changed field. [FileTooLarge ](/docs/reference/errors/file-too-large/)A file request carries at most 64 MiB; larger data does not go through the files API. [Forbidden ](/docs/reference/errors/forbidden/)The credential is valid but lacks the scope or project access the operation needs. [FunctionCallBatchTooLarge ](/docs/reference/errors/function-call-batch-too-large/)A batch create takes at most 1,000 calls and 8 MiB. [FunctionNotFound ](/docs/reference/errors/function-not-found/)The Function does not exist in the project, or it was deleted since the call was prepared. [IdempotencyKeyReused ](/docs/reference/errors/idempotency-key-reused/)The Idempotency-Key was used in the last 24 hours with a different request body. [ImageBuildUnavailable ](/docs/reference/errors/image-build-unavailable/)This deployment cannot build image steps, Dockerfiles or Sandbox snapshots; prebuilt registry images remain supported. [ImageNotFound ](/docs/reference/errors/image-not-found/)The tag or digest does not exist in the registry, or the pull credentials cannot see it. [ImageNotReady ](/docs/reference/errors/image-not-ready/)An imageRef runs the digest the Image built; the Image has none yet. [ImagePullFailed ](/docs/reference/errors/image-pull-failed/)The registry refused the credentials or was unavailable while pinning the image to a digest. [InsufficientCredits ](/docs/reference/errors/insufficient-credits/)The org's available credits do not cover the hold that the action reserves before it starts. [InternalError ](/docs/reference/errors/internal-error/)The server hit an unexpected failure; the request id finds it in the logs. [Invalid ](/docs/reference/errors/invalid/)The object fails strict decoding or validation; \`details.causes\` lists every field and reason. [MethodNotAllowed ](/docs/reference/errors/method-not-allowed/)The resource does not support this verb (for example, a read-only kind). [NoSSHKey ](/docs/reference/errors/no-ssh-key/)The Workspace accepts only the SSHKeys named in spec.sshKeys, or every project member's keys when it names none. [NotFound ](/docs/reference/errors/not-found/)The object, project or path does not exist in this org. [OutputIndexRequired ](/docs/reference/errors/output-index-required/)Several indexes of an Indexed Job committed the output; the download must pick one. [OutputNotFound ](/docs/reference/errors/output-not-found/)The Job did not commit an output by that name, or not at that index. [PaymentDisputed ](/docs/reference/errors/payment-disputed/)A card payment on the org is disputed; new work cannot start until the dispute closes. [PaymentMethodRequired ](/docs/reference/errors/payment-method-required/)New work requires a verified card, including work funded by promotional credit. Adding a card does not charge it. [PaymentsNotConfigured ](/docs/reference/errors/payments-not-configured/)The deployment holds no Stripe key, so it takes no card payments; balances, grants and usage still work. [PaymentVerificationUnavailable ](/docs/reference/errors/payment-verification-unavailable/)Stripe could not confirm the payment method. Existing running work retains its funded holds. [PoolActionNotPending ](/docs/reference/errors/pool-action-not-pending/)Only a queued, undecided action waits for a decision; it may have expired or been decided by someone else. [PoolActionNotReversible ](/docs/reference/errors/pool-action-not-reversible/)Only an executed drain, undrain or routing change can be reverted, once. [PoolActPaused ](/docs/reference/errors/pool-act-paused/)The pool's kill switch stops every action until it is resumed. [PoolInUse ](/docs/reference/errors/pool-in-use/)A Pool is deleted only once no Node is enrolled in it and no object names it in spec.placement.pool. [PreconditionFailed ](/docs/reference/errors/precondition-failed/)The If-Match value no longer matches the compiled spec and price book of a fresh dry-run. [PriceConsentRequired ](/docs/reference/errors/price-consent-required/)Pool routing and Predict are billed, so turning either on records consent to its current price. [PromoCodeExpired ](/docs/reference/errors/promo-code-expired/)The promo code exists but can no longer be redeemed. [PromoCodeUsed ](/docs/reference/errors/promo-code-used/)The org has already redeemed the code as often as it may, or the code has no redemptions left. [QuotaExceeded ](/docs/reference/errors/quota-exceeded/)The org or project already uses everything its Quota allows of this counter. [RequestEntityTooLarge ](/docs/reference/errors/request-entity-too-large/)The request body is larger than the route accepts (1 MiB for JSON). [RequestInProgress ](/docs/reference/errors/request-in-progress/)The original request with this Idempotency-Key has not finished yet. [SandboxFailed ](/docs/reference/errors/sandbox-failed/)The Sandbox is in the Failed phase, which is final: status.reason says why. [SandboxNotRunning ](/docs/reference/errors/sandbox-not-running/)The Sandbox is being deleted, so no command, file or port request can reach it and none starts it. [SandboxStarting ](/docs/reference/errors/sandbox-starting/)The request needs the Sandbox's container, which is not running yet. A stopped Sandbox is started by the request itself; the answer carries Retry-After. [SecretValueInEnv ](/docs/reference/errors/secret-value-in-env/)Secret values reach a Job or Sandbox only through secrets or valueFrom.secretKeyRef, never as literals. [SHA256Mismatch ](/docs/reference/errors/sha256-mismatch/)A write with expectedSHA256 found different content than the writer started from, so nothing was written. [StdinBackpressure ](/docs/reference/errors/stdin-backpressure/)The process's stdin buffer is full, so the write was refused without losing earlier input. [TooManyRequests ](/docs/reference/errors/too-many-requests/)The credential sent too many requests of this class; the RateLimit headers show the budget. [TopUpLimitReached ](/docs/reference/errors/top-up-limit-reached/)An org buys at most $5,000 of credit through Checkout per UTC day. [Unauthorized ](/docs/reference/errors/unauthorized/)The request carries no valid credential, or the credential expired. [Unavailable ](/docs/reference/errors/unavailable/)A dependency of the API is temporarily unavailable. [Unsupported ](/docs/reference/errors/unsupported/)The request is well formed but asks for a feature this API does not offer. [UnsupportedMediaType ](/docs/reference/errors/unsupported-media-type/)The request body's Content-Type is not accepted; strategic merge patch is not supported. [VolumeBusy ](/docs/reference/errors/volume-busy/)A ReadWriteOnce Volume has one writer; another attempt holds its lease. [WorkerNotAllowed ](/docs/reference/errors/worker-not-allowed/)Only the Function's own worker attempts, through their ServiceAccount token, can claim its calls. [WorkspaceStarting ](/docs/reference/errors/workspace-starting/)The Workspace has no running session yet: it is starting, or the request woke it from a stop. [WorkspaceToolUnavailable ](/docs/reference/errors/workspace-tool-unavailable/)A browser tool has a preview URL only while the Workspace is running and lists the tool in spec.tools. # AdmissionsPaused > Operators paused admissions for this kind; existing objects keep running. **HTTP status:** 503 · **Retryable:** yes Operators paused admissions for this kind; existing objects keep running. ## Message [Section titled “Message”](#message) ```text new {resource} are not being admitted right now ``` ## What to do [Section titled “What to do”](#what-to-do) retry later; shows the incident ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: AdmissionsPaused` and `code: 503`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.AdmissionsPaused`. # AgentGroupClosed > A sealed, canceled, finished or deleting AgentGroup admits no new member runs. **HTTP status:** 409 · **Retryable:** no A sealed, canceled, finished or deleting AgentGroup admits no new member runs. ## Message [Section titled “Message”](#message) ```text agent group "{group}" is {state} and takes no new runs ``` ## What to do [Section titled “What to do”](#what-to-do) create the runs in a new group (`nodus create agentgroup --agent `) ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: AgentGroupClosed` and `code: 409`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.AgentGroupClosed`. # AgentGroupFull > An AgentGroup admits at most spec.limits.maxPending unfinished member runs. **HTTP status:** 409 · **Retryable:** yes An AgentGroup admits at most spec.limits.maxPending unfinished member runs. ## Message [Section titled “Message”](#message) ```text agent group "{group}" already has {limit} unfinished runs ``` ## What to do [Section titled “What to do”](#what-to-do) wait for runs to finish, or raise the limit (`nodus patch ag/{group} -p '{"spec":{"limits":{"maxPending":}}}'`) ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: AgentGroupFull` and `code: 409`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.AgentGroupFull`. # AgentRunFinished > A run in a terminal phase (Succeeded, Failed or canceled) no longer reads its inbox. **HTTP status:** 409 · **Retryable:** no A run in a terminal phase (Succeeded, Failed or canceled) no longer reads its inbox. ## Message [Section titled “Message”](#message) ```text agent run "{name}" has finished and takes no more messages ``` ## What to do [Section titled “What to do”](#what-to-do) start a new run with `nodus create agentrun --agent ` ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: AgentRunFinished` and `code: 409`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.AgentRunFinished`. # AgentStopped > An Agent in the Stopped state refuses new runs; runs already created are not affected. **HTTP status:** 409 · **Retryable:** no An Agent in the Stopped state refuses new runs; runs already created are not affected. ## Message [Section titled “Message”](#message) ```text agent "{agent}" is stopped and takes no new runs ``` ## What to do [Section titled “What to do”](#what-to-do) set the agent’s state to Running (`nodus patch agent/{agent} -p '{"spec":{"state":"Running"}}'`) and create the run again ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: AgentStopped` and `code: 409`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.AgentStopped`. # AlreadyExists > A create by name found an object with that name whose spec differs; an identical create returns the object. **HTTP status:** 409 · **Retryable:** no A create by name found an object with that name whose spec differs; an identical create returns the object. ## Message [Section titled “Message”](#message) ```text {resource} "{name}" already exists with a different spec: {fields} differ ``` ## What to do [Section titled “What to do”](#what-to-do) choose another name, delete the existing object, or apply the change with `nodus apply` ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: AlreadyExists` and `code: 409`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.AlreadyExists`. `details` also carries: * `diff`, which the Python SDK exposes as `diff` # ArrearsOutstanding > Charges that the org's credits could not cover are unpaid; they block new work until they are settled. **HTTP status:** 402 · **Retryable:** no Charges that the org’s credits could not cover are unpaid; they block new work until they are settled. ## Message [Section titled “Message”](#message) ```text this org has {arrears} of unpaid charges; new work is paused until they are paid ``` ## What to do [Section titled “What to do”](#what-to-do) add credits with `nodus billing top-up {topUp}`; a top-up pays the unpaid charges first ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: ArrearsOutstanding` and `code: 402`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.ArrearsOutstanding`. `details` also carries: * `arrearsUSD`, which the Python SDK exposes as `arrears_usd` # BadRequest > The request is malformed: a query parameter, selector or body could not be parsed. **HTTP status:** 400 · **Retryable:** no The request is malformed: a query parameter, selector or body could not be parsed. ## Message [Section titled “Message”](#message) ```text {reason} ``` ## What to do [Section titled “What to do”](#what-to-do) correct the request; `nodus explain ` shows the accepted fields ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: BadRequest` and `code: 400`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.BadRequest`. # BudgetExceeded > A Block budget or a maxCostUSD cap that applies to the object has less room than the action's hold. **HTTP status:** 402 · **Retryable:** no A Block budget or a maxCostUSD cap that applies to the object has less room than the action’s hold. ## Message [Section titled “Message”](#message) ```text {limit} has {remaining} left of {limitUSD} {period}; {subject} needs {need} ``` ## What to do [Section titled “What to do”](#what-to-do) raise the limit ({field} on {limitRef}), or wait until {resetsAt} ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: BudgetExceeded` and `code: 402`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.BudgetExceeded`. `details` also carries: * `budget`, which the Python SDK exposes as `budget` * `neededUSD`, which the Python SDK exposes as `needed_usd` * `remainingUSD`, which the Python SDK exposes as `remaining_usd` # CapacityUnavailable > No offering that meets the requirements has capacity at the moment. **HTTP status:** 503 · **Retryable:** no No offering that meets the requirements has capacity at the moment. ## Message [Section titled “Message”](#message) ```text no matching capacity is available right now{where} ``` ## What to do [Section titled “What to do”](#what-to-do) wait, or widen placement.regions or resources.gpu.type ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: CapacityUnavailable` and `code: 503`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.CapacityUnavailable`. `details` also carries: * `eta`, which the Python SDK exposes as `eta` # CloudAccountNotVerified > Nodus reads inventory only after it has verified the account's read-only access. **HTTP status:** 409 · **Retryable:** no Nodus reads inventory only after it has verified the account’s read-only access. ## Message [Section titled “Message”](#message) ```text cloud account "{name}" is not verified: {reason} ``` ## What to do [Section titled “What to do”](#what-to-do) finish onboarding with `nodus get cloudaccount/{name} --subresource onboarding` ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: CloudAccountNotVerified` and `code: 409`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.CloudAccountNotVerified`. # Conflict > The object changed after the caller read it (the resourceVersion is stale). **HTTP status:** 409 · **Retryable:** no The object changed after the caller read it (the resourceVersion is stale). ## Message [Section titled “Message”](#message) ```text {resource} "{name}" was modified since it was read; the update was not applied ``` ## What to do [Section titled “What to do”](#what-to-do) read the object again and reapply the change ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: Conflict` and `code: 409`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.Conflict`. # ConnectionNotReady > A Connection is usable once a verification passed. **HTTP status:** 409 · **Retryable:** yes A Connection is usable once a verification passed. ## Message [Section titled “Message”](#message) ```text connection "{name}" is not verified: {reason} ``` ## What to do [Section titled “What to do”](#what-to-do) fix the credentials Secret and run `nodus describe connection/{name}` to see the last check ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: ConnectionNotReady` and `code: 409`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.ConnectionNotReady`. # EncryptionFailed > The org's encryption key was unavailable; nothing was stored. **HTTP status:** 503 · **Retryable:** yes The org’s encryption key was unavailable; nothing was stored. ## Message [Section titled “Message”](#message) ```text the values of secret "{name}" could not be encrypted ``` ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: EncryptionFailed` and `code: 503`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.EncryptionFailed`. # EnrollmentTokenInvalid > An enrollment token works once, for 24 hours, for the pool it was created for. **HTTP status:** 401 · **Retryable:** no An enrollment token works once, for 24 hours, for the pool it was created for. ## Message [Section titled “Message”](#message) ```text the enrollment token is invalid, expired, revoked or already used ``` ## What to do [Section titled “What to do”](#what-to-do) create a new token with `nodus create enrollmenttoken --pool {pool}` and run the installer it prints ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: EnrollmentTokenInvalid` and `code: 401`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.EnrollmentTokenInvalid`. # Expired > The watch position or list continue token is older than the 24-hour change history. **HTTP status:** 410 · **Retryable:** no The watch position or list continue token is older than the 24-hour change history. ## Message [Section titled “Message”](#message) ```text resourceVersion {resourceVersion} is older than the retained history ``` ## What to do [Section titled “What to do”](#what-to-do) list again and watch from the new resourceVersion ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: Expired` and `code: 410`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.Expired`. # FieldImmutable > The update changes a spec field that is immutable after create; the causes list each changed field. **HTTP status:** 409 · **Retryable:** no The update changes a spec field that is immutable after create; the causes list each changed field. ## Message [Section titled “Message”](#message) ```text {resource} "{name}": {fields} cannot be changed after create ``` ## What to do [Section titled “What to do”](#what-to-do) create a new object with the new values, or revert those fields ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: FieldImmutable` and `code: 409`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.FieldImmutable`. `details` also carries: * `diff`, which the Python SDK exposes as `diff` # FileTooLarge > A file request carries at most 64 MiB; larger data does not go through the files API. **HTTP status:** 413 · **Retryable:** no A file request carries at most 64 MiB; larger data does not go through the files API. ## Message [Section titled “Message”](#message) ```text {path} is larger than {limit} ``` ## What to do [Section titled “What to do”](#what-to-do) put large data on a Volume (`nodus volume put`) and mount it in the Sandbox ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: FileTooLarge` and `code: 413`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.FileTooLarge`. # Forbidden > The credential is valid but lacks the scope or project access the operation needs. **HTTP status:** 403 · **Retryable:** no The credential is valid but lacks the scope or project access the operation needs. ## Message [Section titled “Message”](#message) ```text {principal} cannot {verb} {resource}{where}: missing scope {scope} ``` ## What to do [Section titled “What to do”](#what-to-do) ask an org admin for the {scope} scope, or use a key that has it (`nodus auth can-i {verb} {resource}`) ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: Forbidden` and `code: 403`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.Forbidden`. `details` also carries: * `missingScope`, which the Python SDK exposes as `missing_scope` # FunctionCallBatchTooLarge > A batch create takes at most 1,000 calls and 8 MiB. **HTTP status:** 413 · **Retryable:** no A batch create takes at most 1,000 calls and 8 MiB. ## Message [Section titled “Message”](#message) ```text a FunctionCallList holds at most {max} calls, got {count} ``` ## What to do [Section titled “What to do”](#what-to-do) send `.map()` inputs in batches of at most 1,000 calls; the SDK does this for you ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: FunctionCallBatchTooLarge` and `code: 413`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.FunctionCallBatchTooLarge`. # FunctionNotFound > The Function does not exist in the project, or it was deleted since the call was prepared. **HTTP status:** 404 · **Retryable:** no The Function does not exist in the project, or it was deleted since the call was prepared. ## Message [Section titled “Message”](#message) ```text project "{project}" has no function "{name}" ``` ## What to do [Section titled “What to do”](#what-to-do) list the project’s Functions with `nodus get functions`, or deploy the App with `nodus deploy app.py` ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: FunctionNotFound` and `code: 404`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.FunctionNotFound`. # IdempotencyKeyReused > The Idempotency-Key was used in the last 24 hours with a different request body. **HTTP status:** 409 · **Retryable:** no The Idempotency-Key was used in the last 24 hours with a different request body. ## Message [Section titled “Message”](#message) ```text Idempotency-Key {key} was already used for a different request ``` ## What to do [Section titled “What to do”](#what-to-do) send a new Idempotency-Key for a different request ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: IdempotencyKeyReused` and `code: 409`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.IdempotencyKeyReused`. # ImageBuildUnavailable > This deployment cannot build image steps, Dockerfiles or Sandbox snapshots; prebuilt registry images remain supported. **HTTP status:** 503 · **Retryable:** no This deployment cannot build image steps, Dockerfiles or Sandbox snapshots; prebuilt registry images remain supported. ## Message [Section titled “Message”](#message) ```text image builds are unavailable on this deployment ``` ## What to do [Section titled “What to do”](#what-to-do) publish the image to a registry and use image or a base-only Image without steps; for an Environment use package.image instead of package.pip ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: ImageBuildUnavailable` and `code: 503`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.ImageBuildUnavailable`. # ImageNotFound > The tag or digest does not exist in the registry, or the pull credentials cannot see it. **HTTP status:** 422 · **Retryable:** no The tag or digest does not exist in the registry, or the pull credentials cannot see it. ## Message [Section titled “Message”](#message) ```text image "{image}" was not found ``` ## What to do [Section titled “What to do”](#what-to-do) check the reference, or add a Registry Secret under imagePullSecrets ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: ImageNotFound` and `code: 422`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.ImageNotFound`. # ImageNotReady > An imageRef runs the digest the Image built; the Image has none yet. **HTTP status:** 409 · **Retryable:** yes An imageRef runs the digest the Image built; the Image has none yet. ## Message [Section titled “Message”](#message) ```text image "{name}" is {phase} ``` ## What to do [Section titled “What to do”](#what-to-do) wait with `nodus wait image/{name} --for=condition=Ready`, then retry ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: ImageNotReady` and `code: 409`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.ImageNotReady`. # ImagePullFailed > The registry refused the credentials or was unavailable while pinning the image to a digest. **HTTP status:** 422 · **Retryable:** yes The registry refused the credentials or was unavailable while pinning the image to a digest. ## Message [Section titled “Message”](#message) ```text image "{image}" could not be resolved: {reason} ``` ## What to do [Section titled “What to do”](#what-to-do) check the imagePullSecrets, or retry ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: ImagePullFailed` and `code: 422`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.ImagePullFailed`. # InsufficientCredits > The org's available credits do not cover the hold that the action reserves before it starts. **HTTP status:** 402 · **Retryable:** no The org’s available credits do not cover the hold that the action reserves before it starts. ## Message [Section titled “Message”](#message) ```text {subject} needs a {need} hold to start (released when it ends); available {available} ``` ## What to do [Section titled “What to do”](#what-to-do) add credits with `nodus billing top-up {topUp}`, or lower spec.maxCostUSD ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: InsufficientCredits` and `code: 402`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.InsufficientCredits`. `details` also carries: * `availableUSD`, which the Python SDK exposes as `available_usd` * `neededUSD`, which the Python SDK exposes as `needed_usd` # InternalError > The server hit an unexpected failure; the request id finds it in the logs. **HTTP status:** 500 · **Retryable:** yes The server hit an unexpected failure; the request id finds it in the logs. ## Message [Section titled “Message”](#message) ```text internal error ``` ## What to do [Section titled “What to do”](#what-to-do) retry; if it persists, contact support with the requestId ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: InternalError` and `code: 500`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.InternalError`. # Invalid > The object fails strict decoding or validation; `details.causes` lists every field and reason. **HTTP status:** 422 · **Retryable:** no The object fails strict decoding or validation; `details.causes` lists every field and reason. ## Message [Section titled “Message”](#message) ```text {object} is invalid: {errors} ``` ## What to do [Section titled “What to do”](#what-to-do) fix the listed fields; `nodus explain .` documents each one ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: Invalid` and `code: 422`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.Invalid`. # MethodNotAllowed > The resource does not support this verb (for example, a read-only kind). **HTTP status:** 405 · **Retryable:** no The resource does not support this verb (for example, a read-only kind). ## Message [Section titled “Message”](#message) ```text {verb} is not allowed on {resource} ``` ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: MethodNotAllowed` and `code: 405`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.MethodNotAllowed`. # NoSSHKey > The Workspace accepts only the SSHKeys named in spec.sshKeys, or every project member's keys when it names none. **HTTP status:** 403 · **Retryable:** no The Workspace accepts only the SSHKeys named in spec.sshKeys, or every project member’s keys when it names none. ## Message [Section titled “Message”](#message) ```text no SSH key of yours may log in to workspace "{name}" ``` ## What to do [Section titled “What to do”](#what-to-do) add a key with `nodus create sshkey --from-file ~/.ssh/id_ed25519.pub`, or ask the owner to list yours in spec.sshKeys ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: NoSSHKey` and `code: 403`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.NoSSHKey`. # NotFound > The object, project or path does not exist in this org. **HTTP status:** 404 · **Retryable:** no The object, project or path does not exist in this org. ## Message [Section titled “Message”](#message) ```text {resource} "{name}" not found ``` ## What to do [Section titled “What to do”](#what-to-do) check the name and project (`nodus get {resource} -p `) ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: NotFound` and `code: 404`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.NotFound`. # OutputIndexRequired > Several indexes of an Indexed Job committed the output; the download must pick one. **HTTP status:** 400 · **Retryable:** no Several indexes of an Indexed Job committed the output; the download must pick one. ## Message [Section titled “Message”](#message) ```text output "{output}" of job "{name}" was produced by {count} indexes ``` ## What to do [Section titled “What to do”](#what-to-do) pass ?index=* (`nodus cp job/{name}:{output} . --index `)* *## On the wire[Section titled “On the wire”](#on-the-wire)The response is a Kubernetes `Status` with `reason: OutputIndexRequired` and `code: 400`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.OutputIndexRequired`.* # OutputNotFound > The Job did not commit an output by that name, or not at that index. **HTTP status:** 404 · **Retryable:** no The Job did not commit an output by that name, or not at that index. ## Message [Section titled “Message”](#message) ```text job "{name}" has no committed output "{output}"{where} ``` ## What to do [Section titled “What to do”](#what-to-do) list the committed outputs with `nodus get job/{name} -o jsonpath='{.status.outputs}'` ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: OutputNotFound` and `code: 404`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.OutputNotFound`. # PaymentDisputed > A card payment on the org is disputed; new work cannot start until the dispute closes. **HTTP status:** 402 · **Retryable:** no A card payment on the org is disputed; new work cannot start until the dispute closes. ## Message [Section titled “Message”](#message) ```text a payment on this org is disputed; new work is paused until the dispute closes ``` ## What to do [Section titled “What to do”](#what-to-do) contact support to resolve the dispute ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: PaymentDisputed` and `code: 402`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.PaymentDisputed`. # PaymentMethodRequired > New work requires a verified card, including work funded by promotional credit. Adding a card does not charge it. **HTTP status:** 402 · **Retryable:** no New work requires a verified card, including work funded by promotional credit. Adding a card does not charge it. ## Message [Section titled “Message”](#message) ```text add a valid payment method before starting work ``` ## What to do [Section titled “What to do”](#what-to-do) open Billing and choose Add a card, or run `nodus billing portal --setup`; if you just added it, wait for verification and retry ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: PaymentMethodRequired` and `code: 402`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.PaymentMethodRequired`. # PaymentVerificationUnavailable > Stripe could not confirm the payment method. Existing running work retains its funded holds. **HTTP status:** 503 · **Retryable:** yes Stripe could not confirm the payment method. Existing running work retains its funded holds. ## Message [Section titled “Message”](#message) ```text payment method verification is temporarily unavailable ``` ## What to do [Section titled “What to do”](#what-to-do) retry shortly; contact support if verification remains unavailable ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: PaymentVerificationUnavailable` and `code: 503`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.PaymentVerificationUnavailable`. # PaymentsNotConfigured > The deployment holds no Stripe key, so it takes no card payments; balances, grants and usage still work. **HTTP status:** 503 · **Retryable:** no The deployment holds no Stripe key, so it takes no card payments; balances, grants and usage still work. ## Message [Section titled “Message”](#message) ```text payments are not configured on this deployment ``` ## What to do [Section titled “What to do”](#what-to-do) contact support to restore payment setup and verification ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: PaymentsNotConfigured` and `code: 503`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.PaymentsNotConfigured`. # PoolActPaused > The pool's kill switch stops every action until it is resumed. **HTTP status:** 409 · **Retryable:** no The pool’s kill switch stops every action until it is resumed. ## Message [Section titled “Message”](#message) ```text actions on pool "{pool}" are paused ``` ## What to do [Section titled “What to do”](#what-to-do) resume actions with `nodus pool resume {pool}` ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: PoolActPaused` and `code: 409`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.PoolActPaused`. # PoolActionNotPending > Only a queued, undecided action waits for a decision; it may have expired or been decided by someone else. **HTTP status:** 409 · **Retryable:** no Only a queued, undecided action waits for a decision; it may have expired or been decided by someone else. ## Message [Section titled “Message”](#message) ```text pool action "{name}" is {phase} and no longer waits for a decision ``` ## What to do [Section titled “What to do”](#what-to-do) check it with `nodus get poolaction/{name}` ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: PoolActionNotPending` and `code: 409`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.PoolActionNotPending`. # PoolActionNotReversible > Only an executed drain, undrain or routing change can be reverted, once. **HTTP status:** 409 · **Retryable:** no Only an executed drain, undrain or routing change can be reverted, once. ## Message [Section titled “Message”](#message) ```text pool action "{name}" cannot be reverted: {reason} ``` ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: PoolActionNotReversible` and `code: 409`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.PoolActionNotReversible`. # PoolInUse > A Pool is deleted only once no Node is enrolled in it and no object names it in spec.placement.pool. **HTTP status:** 409 · **Retryable:** no A Pool is deleted only once no Node is enrolled in it and no object names it in spec.placement.pool. ## Message [Section titled “Message”](#message) ```text pool "{pool}" is in use by {users} ``` ## What to do [Section titled “What to do”](#what-to-do) delete its nodes with `nodus delete node ` and move or delete the objects that place on it ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: PoolInUse` and `code: 409`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.PoolInUse`. # PreconditionFailed > The If-Match value no longer matches the compiled spec and price book of a fresh dry-run. **HTTP status:** 412 · **Retryable:** no The If-Match value no longer matches the compiled spec and price book of a fresh dry-run. ## Message [Section titled “Message”](#message) ```text the object or its price changed since the dry-run it was reviewed with{reason} ``` ## What to do [Section titled “What to do”](#what-to-do) run the dry-run again (`--dry-run=server`), review the new estimate and create with its ETag ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: PreconditionFailed` and `code: 412`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.PreconditionFailed`. # PriceConsentRequired > Pool routing and Predict are billed, so turning either on records consent to its current price. **HTTP status:** 422 · **Retryable:** no Pool routing and Predict are billed, so turning either on records consent to its current price. ## Message [Section titled “Message”](#message) ```text {feature} costs ${price}; agree to the price to turn it on ``` ## What to do [Section titled “What to do”](#what-to-do) set the annotation {annotation}: “{price}” on the Pool, or agree in the console ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: PriceConsentRequired` and `code: 422`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.PriceConsentRequired`. `details` also carries: * `annotation`, which the Python SDK exposes as `annotation` * `feature`, which the Python SDK exposes as `feature` * `price`, which the Python SDK exposes as `price` # PromoCodeExpired > The promo code exists but can no longer be redeemed. **HTTP status:** 410 · **Retryable:** no The promo code exists but can no longer be redeemed. ## Message [Section titled “Message”](#message) ```text promo code {code} has expired ``` ## What to do [Section titled “What to do”](#what-to-do) ask whoever gave you the code for a current one, or add credits with `nodus billing top-up` ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: PromoCodeExpired` and `code: 410`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.PromoCodeExpired`. # PromoCodeUsed > The org has already redeemed the code as often as it may, or the code has no redemptions left. **HTTP status:** 409 · **Retryable:** no The org has already redeemed the code as often as it may, or the code has no redemptions left. ## Message [Section titled “Message”](#message) ```text promo code {code} {why} ``` ## What to do [Section titled “What to do”](#what-to-do) see the credit it added with `nodus get creditgrants` ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: PromoCodeUsed` and `code: 409`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.PromoCodeUsed`. # QuotaExceeded > The org or project already uses everything its Quota allows of this counter. **HTTP status:** 429 · **Retryable:** no The org or project already uses everything its Quota allows of this counter. ## Message [Section titled “Message”](#message) ```text quota {quota} exceeded ``` ## What to do [Section titled “What to do”](#what-to-do) delete objects you no longer need, or ask for a higher limit (`nodus get quota default`) ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: QuotaExceeded` and `code: 429`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.QuotaExceeded`. `details` also carries: * `quota`, which the Python SDK exposes as `quota` # RequestEntityTooLarge > The request body is larger than the route accepts (1 MiB for JSON). **HTTP status:** 413 · **Retryable:** no The request body is larger than the route accepts (1 MiB for JSON). ## Message [Section titled “Message”](#message) ```text the request body exceeds {limit} ``` ## What to do [Section titled “What to do”](#what-to-do) upload large content as a blob (`/blobs/v1`) and reference it by sha256 ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: RequestEntityTooLarge` and `code: 413`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.RequestEntityTooLarge`. # RequestInProgress > The original request with this Idempotency-Key has not finished yet. **HTTP status:** 409 · **Retryable:** yes The original request with this Idempotency-Key has not finished yet. ## Message [Section titled “Message”](#message) ```text a request with Idempotency-Key {key} is still in progress ``` ## What to do [Section titled “What to do”](#what-to-do) retry after the Retry-After interval with the same key ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: RequestInProgress` and `code: 409`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.RequestInProgress`. # SandboxFailed > The Sandbox is in the Failed phase, which is final: status.reason says why. **HTTP status:** 409 · **Retryable:** no The Sandbox is in the Failed phase, which is final: status.reason says why. ## Message [Section titled “Message”](#message) ```text sandbox "{name}" failed: {cause} ``` ## What to do [Section titled “What to do”](#what-to-do) read the cause with `nodus describe sandbox/{name}`, then delete the Sandbox and create it again ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: SandboxFailed` and `code: 409`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.SandboxFailed`. # SandboxNotRunning > The Sandbox is being deleted, so no command, file or port request can reach it and none starts it. **HTTP status:** 409 · **Retryable:** no The Sandbox is being deleted, so no command, file or port request can reach it and none starts it. ## Message [Section titled “Message”](#message) ```text sandbox "{name}" is {state} and takes no more requests ``` ## What to do [Section titled “What to do”](#what-to-do) create a new Sandbox once the old one is gone (`nodus get sandbox/{name}` shows when) ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: SandboxNotRunning` and `code: 409`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.SandboxNotRunning`. # SandboxStarting > The request needs the Sandbox's container, which is not running yet. A stopped Sandbox is started by the request itself; the answer carries Retry-After. **HTTP status:** 503 · **Retryable:** yes The request needs the Sandbox’s container, which is not running yet. A stopped Sandbox is started by the request itself; the answer carries Retry-After. ## Message [Section titled “Message”](#message) ```text sandbox "{name}" is starting ``` ## What to do [Section titled “What to do”](#what-to-do) retry after the Retry-After interval, or run `nodus wait sandbox/{name} --for=condition=Ready` ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: SandboxStarting` and `code: 503`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.SandboxStarting`. # SecretValueInEnv > Secret values reach a Job or Sandbox only through secrets or valueFrom.secretKeyRef, never as literals. **HTTP status:** 422 · **Retryable:** no Secret values reach a Job or Sandbox only through secrets or valueFrom.secretKeyRef, never as literals. ## Message [Section titled “Message”](#message) ```text env contains the value of a Secret{where} ``` ## What to do [Section titled “What to do”](#what-to-do) reference the Secret with valueFrom.secretKeyRef or list it under secrets ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: SecretValueInEnv` and `code: 422`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.SecretValueInEnv`. # SHA256Mismatch > A write with expectedSHA256 found different content than the writer started from, so nothing was written. **HTTP status:** 412 · **Retryable:** no A write with expectedSHA256 found different content than the writer started from, so nothing was written. ## Message [Section titled “Message”](#message) ```text the file {path} changed: its SHA-256 does not match expectedSHA256 ``` ## What to do [Section titled “What to do”](#what-to-do) read the file again, merge your change, and write it with the new digest as expectedSHA256 ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: SHA256Mismatch` and `code: 412`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.SHA256Mismatch`. # StdinBackpressure > The process's stdin buffer is full, so the write was refused without losing earlier input. **HTTP status:** 429 · **Retryable:** yes The process’s stdin buffer is full, so the write was refused without losing earlier input. ## Message [Section titled “Message”](#message) ```text the process is not reading stdin as fast as it is sent ``` ## What to do [Section titled “What to do”](#what-to-do) retry after a short backoff; the SDK and CLI do this automatically ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: StdinBackpressure` and `code: 429`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.StdinBackpressure`. # TooManyRequests > The credential sent too many requests of this class; the RateLimit headers show the budget. **HTTP status:** 429 · **Retryable:** yes The credential sent too many requests of this class; the RateLimit headers show the budget. ## Message [Section titled “Message”](#message) ```text rate limit exceeded for {class} requests ``` ## What to do [Section titled “What to do”](#what-to-do) retry after the Retry-After interval ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: TooManyRequests` and `code: 429`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.TooManyRequests`. # TopUpLimitReached > An org buys at most $5,000 of credit through Checkout per UTC day. **HTTP status:** 429 · **Retryable:** no An org buys at most $5,000 of credit through Checkout per UTC day. ## Message [Section titled “Message”](#message) ```text top-ups of this org would exceed {limit} today (UTC) ``` ## What to do [Section titled “What to do”](#what-to-do) buy the rest after midnight UTC, or contact support for a larger purchase ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: TopUpLimitReached` and `code: 429`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.TopUpLimitReached`. # Unauthorized > The request carries no valid credential, or the credential expired. **HTTP status:** 401 · **Retryable:** no The request carries no valid credential, or the credential expired. ## Message [Section titled “Message”](#message) ```text authentication required{reason} ``` ## What to do [Section titled “What to do”](#what-to-do) run `nodus login`, or send a valid API key as `Authorization: Bearer ` ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: Unauthorized` and `code: 401`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.Unauthorized`. # Unavailable > A dependency of the API is temporarily unavailable. **HTTP status:** 503 · **Retryable:** yes A dependency of the API is temporarily unavailable. ## Message [Section titled “Message”](#message) ```text the service is temporarily unavailable{reason} ``` ## What to do [Section titled “What to do”](#what-to-do) retry with backoff ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: Unavailable` and `code: 503`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.Unavailable`. # Unsupported > The request is well formed but asks for a feature this API does not offer. **HTTP status:** 422 · **Retryable:** no The request is well formed but asks for a feature this API does not offer. ## Message [Section titled “Message”](#message) ```text {what} is not supported ``` ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: Unsupported` and `code: 422`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.Unsupported`. # UnsupportedMediaType > The request body's Content-Type is not accepted; strategic merge patch is not supported. **HTTP status:** 415 · **Retryable:** no The request body’s Content-Type is not accepted; strategic merge patch is not supported. ## Message [Section titled “Message”](#message) ```text content type {contentType} is not supported{hint} ``` ## What to do [Section titled “What to do”](#what-to-do) send application/json, application/merge-patch+json, application/json-patch+json or application/apply-patch+yaml ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: UnsupportedMediaType` and `code: 415`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.UnsupportedMediaType`. # VolumeBusy > A ReadWriteOnce Volume has one writer; another attempt holds its lease. **HTTP status:** 409 · **Retryable:** yes A ReadWriteOnce Volume has one writer; another attempt holds its lease. ## Message [Section titled “Message”](#message) ```text volume "{name}" is held by {holder} ``` ## What to do [Section titled “What to do”](#what-to-do) wait for {holder} to finish, or mount the Volume with readOnly: true ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: VolumeBusy` and `code: 409`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.VolumeBusy`. `details` also carries: * `holder`, which the Python SDK exposes as `holder` # WorkerNotAllowed > Only the Function's own worker attempts, through their ServiceAccount token, can claim its calls. **HTTP status:** 403 · **Retryable:** no Only the Function’s own worker attempts, through their ServiceAccount token, can claim its calls. ## Message [Section titled “Message”](#message) ```text this credential is not a worker of function "{name}" ``` ## What to do [Section titled “What to do”](#what-to-do) call the Function with `.remote()`, `.spawn()` or `.map()` instead of claiming calls yourself ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: WorkerNotAllowed` and `code: 403`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.WorkerNotAllowed`. # WorkspaceStarting > The Workspace has no running session yet: it is starting, or the request woke it from a stop. **HTTP status:** 503 · **Retryable:** yes The Workspace has no running session yet: it is starting, or the request woke it from a stop. ## Message [Section titled “Message”](#message) ```text workspace "{name}" is starting; its saved files are being restored ``` ## What to do [Section titled “What to do”](#what-to-do) retry after the Retry-After interval, or run `nodus wait workspace/{name} --for=condition=Ready` ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: WorkspaceStarting` and `code: 503`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.WorkspaceStarting`. # WorkspaceToolUnavailable > A browser tool has a preview URL only while the Workspace is running and lists the tool in spec.tools. **HTTP status:** 409 · **Retryable:** no A browser tool has a preview URL only while the Workspace is running and lists the tool in spec.tools. ## Message [Section titled “Message”](#message) ```text workspace "{name}" serves no {tool} right now ``` ## What to do [Section titled “What to do”](#what-to-do) run `nodus start workspace/{name}` and wait for Ready, or add the tool to spec.tools ## On the wire [Section titled “On the wire”](#on-the-wire) The response is a Kubernetes `Status` with `reason: WorkspaceToolUnavailable` and `code: 409`, plus `fix`, `docs` (this page) and `requestId`. `details.causes` lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises `nodus.errors.WorkspaceToolUnavailable`. # Pricing > Every published Nodus rate, from GPU and CPU list prices to Nodus nodes, storage, egress and models. Pricebook version `2026.10.8`, effective 2026-10-01. All prices are in US dollars. ## How you are billed [Section titled “How you are billed”](#how-you-are-billed) Never above list; your rate is shown before launch and frozen for the run. You pay what the provider bills for your machine, from the moment it is created until it is deleted, at provider cost ÷ 0.875 (Nodus keeps 12.5 % of what you pay). Each charge is itemized as `Boot`, `Running`, `Restore` or `Teardown`. Nodus-operated capacity bills at the node rates below. Credit is prepaid and every paid action reserves a hold first. ## GPUs [Section titled “GPUs”](#gpus) List price per hour for the whole machine, by GPU count. | GPU | Memory | ×1 | ×2 | ×4 | ×8 | | ------------- | ------ | ----- | ------ | ------ | ------ | | A10 | 24 GB | $0.89 | $1.78 | $3.56 | $7.12 | | A100 40G | 40 GB | $1.49 | $2.98 | $5.96 | $11.92 | | A100 40G PCIE | 40 GB | $1.39 | $2.78 | $5.56 | $11.12 | | A100 80G | 80 GB | $1.99 | $3.98 | $7.96 | $15.92 | | A100 80G PCIE | 80 GB | $1.89 | $3.78 | $7.56 | $15.12 | | B200 | 180 GB | $6.49 | $12.98 | $25.96 | $51.92 | | H100 PCIE | 80 GB | $2.99 | $5.98 | $11.96 | $23.92 | | H100 SXM | 80 GB | $3.29 | $6.58 | $13.16 | $26.32 | | H200 | 141 GB | $4.29 | $8.58 | $17.16 | $34.32 | | L4 | 24 GB | $0.89 | $1.78 | $3.56 | $7.12 | | L40S | 48 GB | $1.29 | $2.58 | $5.16 | $10.32 | | RTX 3090 | 24 GB | $0.59 | $1.18 | $2.36 | $4.72 | | RTX 4090 | 24 GB | $0.69 | $1.38 | $2.76 | $5.52 | | RTX 6000 ADA | 48 GB | $1.09 | $2.18 | $4.36 | $8.72 | | RTX A6000 | 48 GB | $0.79 | $1.58 | $3.16 | $6.32 | ## CPU machines [Section titled “CPU machines”](#cpu-machines) | Shape | vCPU | Memory | Disk | Per hour | | ------------ | ---- | ------- | ---- | -------- | | `cpu-2x4` | 2 | 4 GiB | — | $0.16 | | `cpu-2x8` | 2 | 8 GiB | — | $0.12 | | `cpu-4x8` | 4 | 8 GiB | — | $0.31 | | `cpu-4x16` | 4 | 16 GiB | — | $0.23 | | `cpu-8x16` | 8 | 16 GiB | — | $0.62 | | `cpu-8x32` | 8 | 32 GiB | — | $0.45 | | `cpu-16x32` | 16 | 32 GiB | — | $0.79 | | `cpu-16x64` | 16 | 64 GiB | — | $0.89 | | `cpu-32x64` | 32 | 64 GiB | — | $1.58 | | `cpu-32x128` | 32 | 128 GiB | — | $1.78 | | `cpu-64x256` | 64 | 256 GiB | — | $3.56 | ## Nodus nodes [Section titled “Nodus nodes”](#nodus-nodes) Sandboxes, CPU Functions, CPU Workspaces and warm pools run on Nodus-operated nodes, billed per second from these rates. | Resource | Per hour | | --------------------------------- | -------- | | vCPU | $0.048 | | GiB of memory | $0.006 | | GiB of disk above 10 GiB per vCPU | $0.00014 | | Example | Per hour | | ----------------------- | -------- | | Smallest Sandbox | $0.015 | | 2 vCPU / 4 GiB Sandbox | $0.12 | | 8 vCPU / 32 GiB CPU Job | $0.576 | ## Storage and egress [Section titled “Storage and egress”](#storage-and-egress) | Meter | Included | Price above the inclusion | | -------------------- | ---------------------- | ------------------------- | | Storage | 10 GB per org | $0.017143 per GB-month | | Egress | 10 GiB per org per day | $0.09 per GiB | | Relayed gang traffic | — | $0.011429 per GiB | ## Bring your own machines [Section titled “Bring your own machines”](#bring-your-own-machines) | Item | Price | | --------------------------------------------- | --------------------- | | Route: Nodus-scheduled work on your pool GPUs | $0.02 per device-hour | | Predict | $99.00 per pool-month | ## Models [Section titled “Models”](#models) Hosted inference is billed per request from its usage. | Model | Operation | Price | | --------------------------------- | -------------- | ---------------------------- | | `baai/bge-m3` | input | $0.013685 per 1M tokens | | `canopylabs/orpheus-v1-english` | speech | $23.157895 per 1M characters | | `deepseek-ai/deepseek-v4-flash` | input | $0.052632 per 1M tokens | | `deepseek-ai/deepseek-v4-flash` | cacheRead | $0.046211 per 1M tokens | | `deepseek-ai/deepseek-v4-flash` | output | $0.305264 per 1M tokens | | `deepseek-ai/deepseek-v4-pro` | input | $0.410527 per 1M tokens | | `deepseek-ai/deepseek-v4-pro` | cacheRead | $0.332527 per 1M tokens | | `deepseek-ai/deepseek-v4-pro` | output | $3.673685 per 1M tokens | | `deepseek-ai/deepseek-v4.1-flash` | input | $0.105264 per 1M tokens | | `deepseek-ai/deepseek-v4.1-flash` | cacheRead | $0.047369 per 1M tokens | | `deepseek-ai/deepseek-v4.1-flash` | output | $0.463158 per 1M tokens | | `moonshotai/kimi-k3` | input | $1.51579 per 1M tokens | | `moonshotai/kimi-k3` | cacheRead | $0.31579 per 1M tokens | | `moonshotai/kimi-k3` | output | $9.473685 per 1M tokens | | `openai/gpt-oss-120b` | input | $0.157895 per 1M tokens | | `openai/gpt-oss-120b` | cacheRead | $0.078948 per 1M tokens | | `openai/gpt-oss-120b` | output | $0.631579 per 1M tokens | | `openai/gpt-oss-20b` | input | $0.052632 per 1M tokens | | `openai/gpt-oss-20b` | cacheRead | $0.005264 per 1M tokens | | `openai/gpt-oss-20b` | output | $0.210527 per 1M tokens | | `openai/whisper-large-v3` | audio | $0.001948 per audio minute | | `openai/whisper-large-v3-turbo` | audio | $0.000702 per audio minute | | `qwen/qwen3.8-27b` | input | $0.442106 per 1M tokens | | `qwen/qwen3.8-27b` | cacheRead | $0.089474 per 1M tokens | | `qwen/qwen3.8-27b` | output | $3.157895 per 1M tokens | | `zai-org/glm-5.2` | input | $0.242106 per 1M tokens | | `zai-org/glm-5.2` | cacheRead | $0.196948 per 1M tokens | | `zai-org/glm-5.2` | output | $4.631579 per 1M tokens | | `zai-org/glm-5.3` | input | $0.231579 per 1M tokens | | `zai-org/glm-5.3` | cacheRead | $0.186843 per 1M tokens | | `zai-org/glm-5.3` | output | $3.568422 per 1M tokens | | `zai-org/glm-5.3-flash` | input | $0.205264 per 1M tokens | | `zai-org/glm-5.3-flash` | cacheRead | $0.041053 per 1M tokens | | `zai-org/glm-5.3-flash` | output | $0.684211 per 1M tokens | | `nodus/indra` | routing input | $0.044211 per 1M tokens | | `nodus/indra` | routing output | $0.00 per 1M tokens | ## Plans [Section titled “Plans”](#plans) A plan’s allowance pays for its models at the prices above; it expires at the end of each month, and usage beyond it draws from your credits. | Plan | Per month | Allowance | Models | | -------- | --------- | --------- | ------------- | | composer | $20.00 | $20.00 | `nodus/indra` | ## What you pay for [Section titled “What you pay for”](#what-you-pay-for) | Item | Who pays | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------ | | Your rented machine for every second the provider bills it, from creation to confirmed deletion: boot, image pull, restore, running, teardown and the provider’s rounding | You, at provider cost ÷ 0.875, never above list | | Sandboxes, Functions, builds and CPU work on Nodus nodes, from placement to release | You, at the published node rates | | Model tokens, including the routing call of nodus/auto | You, at the cheapest available model cost ÷ 0.95 | | Agent runs on Nodus-managed models | You, at model cost ÷ 0.875, from credits | | Storage above the included amount and egress above the daily inclusion | You, at the published rates | | Spare machines Nodus starts to finish sooner, and failures Nodus causes | Nodus | | Requests whose outcome Nodus cannot confirm | Nodus | | Logs, and egress within the daily inclusion | Nodus | ## Worked examples [Section titled “Worked examples”](#worked-examples) | Example | Cost | | --------------------------------------- | --------- | | One H100 for one hour, at most | $3.29 | | A 2 vCPU / 4 GiB Sandbox for one hour | $0.12 | | 1M output tokens on openai/gpt-oss-120b | $0.631579 | ## Starter grant [Section titled “Starter grant”](#starter-grant) The first org a verified user creates receives $30 of credit that expires 30 days after it is granted. Grant credit is spent before purchased credit. # Python SDK reference > Every public module, class and function of the nodus-compute package, generated from its docstrings. Install the SDK with `pip install nodus-compute` and import it as `nodus` (Python 3.10 or newer). The pages below are generated from the package’s docstrings and type annotations. [nodus ](/docs/reference/python/nodus/)Nodus Python SDK (\`pip install nodus-compute\`). [nodus.agent ](/docs/reference/python/nodus-agent/)Agents: durable runs on the journal engine (resources.md §4.1 to §4.3, ADR-048, ADR-101). [nodus.api ](/docs/reference/python/nodus-api/)Layer 1 of the SDK: the generic resource client over every kind (resources.md §10.3). [nodus.api.models ](/docs/reference/python/nodus-api-models/)Pydantic v2 models for every kind, generated from \`api/openapi\` by \`make gen\` (datamodel-code-generator). [nodus.app ](/docs/reference/python/nodus-app/)\`App\`: the group of Functions defined in one Python file (resources.md §3.9, ADR-047). [nodus.checkpoint ](/docs/reference/python/nodus-checkpoint/)The checkpoint handshake and gang checkpoints (resources.md §8, distributed-training.md §10.3). [nodus.cluster ](/docs/reference/python/nodus-cluster/)\`nodus.cluster.info()\`: the gang env contract of distributed-training.md §3.6, as one object. [nodus.envs ](/docs/reference/python/nodus-envs/)Helpers for building your own Environment and datasets: \`nodus.envs.split\`. [nodus.errors ](/docs/reference/python/nodus-errors/)Errors raised by the Nodus SDK. [nodus.examples ](/docs/reference/python/nodus-examples/)Runnable examples that ship with the SDK: \`python -m nodus.examples.rl\` trains a model with RL. [nodus.examples.rl ](/docs/reference/python/nodus-examples-rl/)Reinforcement learning in one file: teach a small model to spell words backwards. [nodus.functions ](/docs/reference/python/nodus-functions/)Functions, FunctionCalls and \`@app.cls\` classes (resources.md §3.10, §3.11, ADR-047, ADR-110). [nodus.image ](/docs/reference/python/nodus-image/)The Modal-shaped \`Image\` builder (resources.md §5.3). [nodus.integrations ](/docs/reference/python/nodus-integrations/)Framework callbacks that report training progress to Nodus from inside a container. [nodus.integrations.lightning ](/docs/reference/python/nodus-integrations-lightning/)\`NodusCallback\` for a Lightning \`Trainer\` (ADR-103). [nodus.integrations.transformers ](/docs/reference/python/nodus-integrations-transformers/)\`NodusCallback\` for the Hugging Face \`Trainer\` and every TRL trainer built on it (ADR-103). [nodus.job ](/docs/reference/python/nodus-job/)\`Job\`: a run-to-completion container with checkpoints, recovery and outputs (resources.md §3.1). [nodus.llm ](/docs/reference/python/nodus-llm/)OpenAI- and Anthropic-compatible inference (resources.md §4.5, §4.6, ADR-049, ADR-100). [nodus.log ](/docs/reference/python/nodus-log/)Structured telemetry from inside a container: metrics, RL task outcomes, work units and phases. [nodus.outputs ](/docs/reference/python/nodus-outputs/)\`nodus.outputs.verify()\`: fail fast inside a Job when its model output would not load (resources.md §10.5). [nodus.process ](/docs/reference/python/nodus-process/)Processes and files inside a running Job, Sandbox or Workspace (resources.md §3.6). [nodus.recipes ](/docs/reference/python/nodus-recipes/)Training recipes: TrainingJob builders for fine-tuning, pretraining, distillation, preference training and RL. [nodus.recipes.finetune ](/docs/reference/python/nodus-recipes-finetune/)Fine-tuning, preference training, distillation and pretraining recipes on the catalog runtimes. [nodus.recipes.rl ](/docs/reference/python/nodus-recipes-rl/)Reinforcement learning and evaluation on catalog Environments or your own. [nodus.runtime ](/docs/reference/python/nodus-runtime/)In-container helpers that speak nodusd's sockets (resources.md §8); each is a no-op outside Nodus. [nodus.sandbox ](/docs/reference/python/nodus-sandbox/)\`Sandbox\`: an isolated long-running container driven by exec, files and tunnels (resources.md §3.5, ADR-044). [nodus.secret ](/docs/reference/python/nodus-secret/)\`Secret\` references (resources.md §5.2). Values are write-only: nothing the API returns contains them. [nodus.sweep ](/docs/reference/python/nodus-sweep/)\`Sweep\`: one Job template run across a matrix of GPUs, regions and parameters (resources.md §3.3, ADR-051). [nodus.volume ](/docs/reference/python/nodus-volume/)\`Volume\`: named storage with Modal's commit and reload semantics for \`ReadWriteMany\` (resources.md §5.1, ADR-091). [nodus.workspace ](/docs/reference/python/nodus-workspace/)\`Workspace\`: a development machine with SSH, VS Code, JupyterLab and a persistent home (resources.md §3.8). # nodus > Nodus Python SDK (`pip install nodus-compute`). Nodus Python SDK (`pip install nodus-compute`). Layer 2, the Modal-shaped layer, is what most code uses: `App`, `@app.function` with `.remote()`, `.spawn()` and `.map()`, `@app.cls`, `Image`, `Volume`, `Secret`, `Sandbox`, `Job`, `Sweep`, `Workspace`, `Agent`, `nodus.llm` and `@nodus.clustered`. Layer 1, `nodus.api`, is the generic resource client under it. Every blocking call also has an `.aio` form for asyncio code (`await f.remote.aio(x)`). Errors are in `nodus.errors`. Submodules load on first attribute access, so `import nodus` stays fast and in-container helpers such as `nodus.progress` pull in nothing they do not need. ## Exports [Section titled “Exports”](#exports) * `Agent`: defined in `nodus.agent` * `AgentGroup`: defined in `nodus.agent` * `AgentRun`: defined in `nodus.agent` * `App`: defined in `nodus.app` * `ClaudeAgent`: defined in `nodus._runtime.agents.claude` * `Client`: defined in `nodus.api` * `Cls`: defined in `nodus.functions` * `Distributed`: defined in `nodus.job` * `Egress`: defined in `nodus._spec` * `Function`: defined in `nodus.functions` * `FunctionCall`: defined in `nodus.functions` * `GPU`: defined in `nodus._spec` * `Image`: defined in `nodus.image` * `InferenceEndpoint`: defined in `nodus.llm` * `Init`: defined in `nodus.sandbox` * `Job`: defined in `nodus.job` * `Process`: defined in `nodus.process` * `Retries`: defined in `nodus._spec` * `RunContext`: defined in `nodus.agent` * `Sandbox`: defined in `nodus.sandbox` * `Secret`: defined in `nodus.secret` * `Service`: defined in `nodus.sandbox` * `Sweep`: defined in `nodus.sweep` * `TrainingJob`: defined in `nodus.recipes` * `Tunnel`: defined in `nodus.sandbox` * `Volume`: defined in `nodus.volume` * `Workspace`: defined in `nodus.workspace` * `clustered`: defined in `nodus.functions` * `enter`: defined in `nodus.functions` * `exit`: defined in `nodus.functions` * `method`: defined in `nodus.functions` * `progress`: defined in `nodus.runtime` * `restored`: defined in `nodus.runtime` * `self`: defined in `nodus.runtime` * `state_dir`: defined in `nodus.runtime` # nodus.agent > Agents: durable runs on the journal engine (resources.md §4.1 to §4.3, ADR-048, ADR-101). Agents: durable runs on the journal engine (resources.md §4.1 to §4.3, ADR-048, ADR-101). This module is the client side: defining and deploying an Agent, submitting AgentRuns, waiting for results, and fanning out through AgentGroups. Inside a run, the agent runtime (`nodus._runtime.agents`) executes the entrypoint with a `RunContext` and journals every `@agent.step`; outside a run a step is a plain function call. The API serves an Agent’s `image` and `perRunMaxCostUSD`, an AgentRun’s `input`, `deadline` and group membership (`group`, `taskKey`, `dependsOn`), a run’s `answer`, `steps`, `messages` and `cancel`, and an AgentGroup’s `agent`, `limits`, `maxCostUSD`, `sealed`, `state` and bounded Environment `evaluation` (ADR-119 defers the rest). Anything else a call names raises `errors.Unsupported` before anything is sent, because the API rejects an unknown field outright; `nodus.ClaudeAgent` defines an agent on the served kind. ## `Agent` [Section titled “Agent”](#agent) ```python class Agent(name: str, *, image: Any = None, source: str | Path | dict[str, Any] | None = None, entrypoint: str | None = None, setup: str | None = None, secrets: list[Any] | None = None, env: dict[str, str] | None = None, network: _spec.Egress | None = None, models: list[str] | None = None, min_workers: int | None = None, max_workers: int | None = None, scaledown_window: Any = None, per_run_max_cost: Any = None, max_cost: Any = None, cpu: Any = None, memory: Any = None, app: Any = None, project: str | None = None) -> None ``` A durable agent definition; deploy it, then submit runs with `.remote()`, `.spawn()` or `.map()`. ### `Agent.deploy` [Section titled “Agent.deploy”](#agentdeploy) ```python deploy() -> View ``` Create or update the Agent; every accepted change is a new revision, and new runs pin it. ### `Agent.entrypoint` [Section titled “Agent.entrypoint”](#agententrypoint) ```python entrypoint(fn: Callable[..., Any]) -> Callable[..., Any] ``` Mark `fn(ctx, input)` as the run entrypoint; its module is uploaded as the Agent’s source. ### `Agent.from_name` [Section titled “Agent.from_name”](#agentfrom_name) ```python from_name(name: str, project: str | None = None) -> _Agent ``` A deployed Agent, submitted to without redeploying it. ### `Agent.local` [Section titled “Agent.local”](#agentlocal) ```python local(input: Any = None) -> Any ``` Run the entrypoint in this process with an in-memory journal (needs the agent runtime). ### `Agent.map` [Section titled “Agent.map”](#agentmap) ```python map(inputs: Iterable[Any], *, max_active: int | None = None, order_outputs: bool = True, return_exceptions: bool = False) -> AsyncIterator[Any] ``` One run per input in an ephemeral AgentGroup; yields each run’s answer text as the run finishes. Answers come in input order unless `order_outputs=False`. A run that does not succeed raises `AgentRunFailed`, or is yielded as that exception with `return_exceptions=True`. ### `Agent.name` [Section titled “Agent.name”](#agentname) Type: `str` ### `Agent.remote` [Section titled “Agent.remote”](#agentremote) ```python remote(input: Any = None, **kwargs: Any) -> Any ``` Submit, wait and return the run’s answer text; raises `AgentRunFailed`. ### `Agent.spawn` [Section titled “Agent.spawn”](#agentspawn) ```python spawn(input: Any = None, **kwargs: Any) -> _AgentRun ``` The same as `submit`. ### `Agent.step` [Section titled “Agent.step”](#agentstep) ```python step(_fn: Callable[..., Any] | None = None, *, effect: str = 'pure', name: str | None = None) -> Any ``` Journal a function as a step: `pure`, `idempotent` or `external` (an unknown outcome needs resolution). ### `Agent.submit` [Section titled “Agent.submit”](#agentsubmit) ```python submit(input: Any = None, *, idempotency_key: str | None = None, session_key: str | None = None, deadline: str | None = None, group: str | None = None, hold_worker: str | None = None, name: str | None = None) -> _AgentRun ``` Start a run and return its handle. A `name` derived from an event makes redeliveries idempotent. `group` is refused (submit group runs through `AgentGroup.submit_many`), and so are `session_key` and a `hold_worker` other than `"auto"`, which the API does not serve yet. ## `AgentGroup` [Section titled “AgentGroup”](#agentgroup) ```python class AgentGroup(obj: Obj) -> None ``` Runs of one Agent under a shared concurrency limit and cost cap, with task dependencies and one cancel. ### `AgentGroup.cancel` [Section titled “AgentGroup.cancel”](#agentgroupcancel) ```python cancel() -> None ``` Cancel every unfinished run in the group. ### `AgentGroup.create` [Section titled “AgentGroup.create”](#agentgroupcreate) ```python create(name: str, agent: _Agent | str, *, max_active: int | None = None, max_pending: int | None = None, max_held: int | None = None, max_cost: Any = None, evaluation: dict[str, Any] | None = None, project: str | None = None) -> _AgentGroup ``` Create the group; `max_active` runs go at once and runs are released while `max_cost` has room for them. `max_cost` caps the runs’ Claude usage (model and routing calls) together; their sandboxes are billed apart. `evaluation={"environment": "nodus/arithmetic-v2@2.0.0", "tasks": 10, "seed": 42}` generates and scores a fixed batch automatically; split defaults to test, repetitions to 1, and tasks times repetitions is at most 100. `timeout` defaults to “30m” (1m to 24h) from group creation, including queued time. Evaluation and agent Sandbox compute is billed separately from the model-only `max_cost`; use a project Budget to cap total spend. `max_held` remains unsupported. ### `AgentGroup.delete` [Section titled “AgentGroup.delete”](#agentgroupdelete) ```python delete() -> None ``` Delete the group and, with it, its runs. ### `AgentGroup.from_name` [Section titled “AgentGroup.from_name”](#agentgroupfrom_name) ```python from_name(name: str, project: str | None = None) -> _AgentGroup ``` A group that exists already, such as one created from YAML. ### `AgentGroup.name` [Section titled “AgentGroup.name”](#agentgroupname) Type: `str` ### `AgentGroup.results` [Section titled “AgentGroup.results”](#agentgroupresults) ```python results() -> list[View] ``` Per-case evaluation outcomes and grader evidence; no hidden answers or task payloads. ### `AgentGroup.runs` [Section titled “AgentGroup.runs”](#agentgroupruns) ```python runs() -> list[_AgentRun] ``` The group’s member runs. ### `AgentGroup.seal` [Section titled “AgentGroup.seal”](#agentgroupseal) ```python seal() -> None ``` Close the group to new runs; it finishes once every run is terminal. ### `AgentGroup.submit_many` [Section titled “AgentGroup.submit_many”](#agentgroupsubmit_many) ```python submit_many(tasks: list[dict[str, Any]]) -> list[_AgentRun] ``` Create one run per task `{key, input, depends_on?}` and return them in the caller’s order. A task starts after the tasks in `depends_on`, which name tasks of this call or of an earlier one. Runs go out in `AgentRunList` batches (each batch is all-or-nothing); a cycle or a repeated key raises `Invalid`. ### `AgentGroup.wait` [Section titled “AgentGroup.wait”](#agentgroupwait) ```python wait(timeout: float | None = None) -> View ``` Return the terminal status; submitted batches must be sealed, while evaluations close automatically. ## `AgentRun` [Section titled “AgentRun”](#agentrun) ```python class AgentRun(obj: Obj) -> None ``` One AgentRun: `wait()`, `result()` (the answer text), `answer()`, `steps()`, `send()` and `cancel()`. ### `AgentRun.answer` [Section titled “AgentRun.answer”](#agentrunanswer) ```python answer() -> str ``` The run’s full answer text. ### `AgentRun.cancel` [Section titled “AgentRun.cancel”](#agentruncancel) ```python cancel() -> None ``` ### `AgentRun.children` [Section titled “AgentRun.children”](#agentrunchildren) ```python children() -> list[_AgentRun] ``` ### `AgentRun.from_name` [Section titled “AgentRun.from_name”](#agentrunfrom_name) ```python from_name(name: str, project: str | None = None) -> _AgentRun ``` ### `AgentRun.logs` [Section titled “AgentRun.logs”](#agentrunlogs) ```python logs(follow: bool = False) -> AsyncIterator[str] ``` ### `AgentRun.name` [Section titled “AgentRun.name”](#agentrunname) Type: `str` ### `AgentRun.outputs` [Section titled “AgentRun.outputs”](#agentrunoutputs) Type: `_Outputs` ### `AgentRun.resolve` [Section titled “AgentRun.resolve”](#agentrunresolve) ```python resolve(step_id: str, decision: str, evidence: dict[str, str] | None = None, result: Any = None, checkpoint_seq: int | None = None, expected_revision: int | None = None) -> View ``` Resolve an `external` step with an unknown outcome: `Completed`, `NoEffect` or `Cancelled`. ### `AgentRun.result` [Section titled “AgentRun.result”](#agentrunresult) ```python result() -> str ``` The run’s answer text (waits for the run first); raises `AgentRunFailed` unless it succeeded. ### `AgentRun.resume` [Section titled “AgentRun.resume”](#agentrunresume) ```python resume() -> None ``` ### `AgentRun.retry` [Section titled “AgentRun.retry”](#agentrunretry) ```python retry() -> None ``` ### `AgentRun.send` [Section titled “AgentRun.send”](#agentrunsend) ```python send(name: str, payload: Any, message_key: str | None = None) -> View ``` ### `AgentRun.steps` [Section titled “AgentRun.steps”](#agentrunsteps) ```python steps() -> list[View] ``` ### `AgentRun.suspend` [Section titled “AgentRun.suspend”](#agentrunsuspend) ```python suspend() -> None ``` ### `AgentRun.wait` [Section titled “AgentRun.wait”](#agentrunwait) ```python wait(timeout: float | None = None) -> View ``` Block until the run is terminal; raises `AgentRunFailed` unless it succeeded. ## `RunContext` [Section titled “RunContext”](#runcontext) ```python class RunContext(Protocol) ``` What an agent entrypoint receives as `ctx`; the agent runtime provides the implementation. ### `RunContext.child_output` [Section titled “RunContext.child_output”](#runcontextchild_output) ```python child_output(child: Any, name: str) -> Any ``` ### `RunContext.continue_as_new` [Section titled “RunContext.continue_as_new”](#runcontextcontinue_as_new) ```python continue_as_new(input: Any) -> None ``` ### `RunContext.gather` [Section titled “RunContext.gather”](#runcontextgather) ```python gather(children: list[Any], return_exceptions: bool = False) -> list[Any] ``` ### `RunContext.idempotency_key` [Section titled “RunContext.idempotency_key”](#runcontextidempotency_key) Type: `str` ### `RunContext.map` [Section titled “RunContext.map”](#runcontextmap) ```python map(inputs: Iterable[Any]) -> list[Any] ``` ### `RunContext.save_output` [Section titled “RunContext.save_output”](#runcontextsave_output) ```python save_output(name: str, path: str) -> None ``` ### `RunContext.send` [Section titled “RunContext.send”](#runcontextsend) ```python send(target: str, name: str, payload: Any) -> None ``` ### `RunContext.sleep` [Section titled “RunContext.sleep”](#runcontextsleep) ```python sleep(duration: str | float) -> None ``` ### `RunContext.sleep_until` [Section titled “RunContext.sleep_until”](#runcontextsleep_until) ```python sleep_until(time: str) -> None ``` ### `RunContext.spawn` [Section titled “RunContext.spawn”](#runcontextspawn) ```python spawn(input: Any, key: str, permissions: Any = None, deadline: str | None = None) -> Any ``` ### `RunContext.state_dir` [Section titled “RunContext.state_dir”](#runcontextstate_dir) Type: `Path` ### `RunContext.step` [Section titled “RunContext.step”](#runcontextstep) ```python step(name: str, fn: Callable[[], Any], effect: str = 'pure') -> Any ``` ### `RunContext.wait_for_message` [Section titled “RunContext.wait_for_message”](#runcontextwait_for_message) ```python wait_for_message(name: str, timeout: str | float | None = None) -> Any ``` # nodus.api > Layer 1 of the SDK: the generic resource client over every kind (resources.md §10.3). Layer 1 of the SDK: the generic resource client over every kind (resources.md §10.3). `nodus.api.get`, `list`, `watch`, `apply`, `create`, `patch`, `delete`, `estimate`, `logs`, `exec`, `files`, `healthz` and `readyz` use the default client; `nodus.Client(...)` makes another one. Every call blocks and has an `.aio` form. Objects are JSON dicts shaped like the kinds in resources.md; `nodus.api.models` holds the generated pydantic models. ## Exports [Section titled “Exports”](#exports) * `Config`: defined in `nodus.api._config` * `KINDS`: defined in `nodus.api._kinds` * `Resource`: defined in `nodus.api._kinds` * `default_client`: defined in `nodus.api._client` * `resolve`: defined in `nodus.api._config` # nodus.api.models > Pydantic v2 models for every kind, generated from `api/openapi` by `make gen` (datamodel-code-generator). Pydantic v2 models for every kind, generated from `api/openapi` by `make gen` (datamodel-code-generator). The `nodus.dev/v1` kinds and their parts are at the top level (`nodus.api.models.Job`); `v1beta1` holds the `nodus.dev/v1beta1` ones (`nodus.api.models.v1beta1.TrainingJob`). The SDK’s calls exchange plain JSON dicts, so a model is opt-in: `Job.model_validate(obj)` reads an object and `job.model_dump(mode="json", by_alias=True, exclude_none=True)` writes it back. Fields a newer server adds are kept. Every module here but this one is generated; never edit them by hand. # nodus.app > `App`: the group of Functions defined in one Python file (resources.md §3.9, ADR-047). `App`: the group of Functions defined in one Python file (resources.md §3.9, ADR-047). `with app.run():` creates an ephemeral App (renewed every 30 s, deleted on exit, so a vanished client leaves nothing running); `app.deploy()` creates or updates the persistent App and prunes members removed from the file. Code is uploaded once per content hash. Decorating never touches the network, so the worker can import the same file. ## `App` [Section titled “App”](#app) ```python class App(name: str, *, project: str | None = None, image: Any = None, secrets: list[Any] | None = None, volumes: dict[str, Any] | None = None, description: str | None = None) -> None ``` ### `App.cls` [Section titled “App.cls”](#appcls) ```python cls(_cls: type | None = None, **options: Any) -> Any ``` Register a class as one Function; `@nodus.method()`s are callable remotely, hooks run once per worker. ### `App.deploy` [Section titled “App.deploy”](#appdeploy) ```python deploy() -> View ``` Create or update the persistent App and its members; members no longer in the file are deleted. ### `App.function` [Section titled “App.function”](#appfunction) ```python function(_fn: Callable[..., Any] | None = None, *, name: str | None = None, gpu: Any = None, cpu: Any = None, memory: Any = None, ephemeral_disk: Any = None, image: Any = None, secrets: list[Any] | None = None, volumes: dict[str, Any] | None = None, env: dict[str, str] | None = None, network: Any = None, timeout: Any = None, retries: Any = None, checkpoint: Any = None, interruptible: Any = None, region: Any = None, profile: Any = None, max_cost: Any = None, min_workers: int | None = None, max_workers: int | None = None, scaledown_window: Any = None, target_concurrency: int | None = None, min_containers: int | None = None, max_containers: int | None = None) -> Any ``` Register a Function; the arguments map onto Function spec fields (resources.md §10.4). ### `App.local_entrypoint` [Section titled “App.local_entrypoint”](#applocal_entrypoint) ```python local_entrypoint(_fn: Callable[..., Any] | None = None, *, name: str | None = None) -> Any ``` Mark the function `nodus run app.py` calls inside an ephemeral run of this App. ### `App.name` [Section titled “App.name”](#appname) Type: `str` ### `App.registered_entrypoints` [Section titled “App.registered_entrypoints”](#appregistered_entrypoints) Type: `dict[str, Callable[..., Any]]` ### `App.registered_functions` [Section titled “App.registered_functions”](#appregistered_functions) Type: `dict[str, _Function]` ### `App.run` [Section titled “App.run”](#apprun) ```python run(*, show_logs: bool = True) -> AsyncIterator[_App] ``` Run the App ephemerally for the duration of the `with` block; it is deleted on exit. # nodus.checkpoint > The checkpoint handshake and gang checkpoints (resources.md §8, distributed-training.md §10.3). The checkpoint handshake and gang checkpoints (resources.md §8, distributed-training.md §10.3). Nodus owns checkpoint cadence and storage; programs write their own files into `nodus.state_dir()`. Before a snapshot the node sends `checkpoint.request`: `on_request(callback)` saves, then `ack()` tells the node the files are consistent. Restoring files never restores process memory: a restarted program loads what it saved. ## `ack` [Section titled “ack”](#ack) ```python ack(seq: str | None = None) -> None ``` Tell the node the checkpoint files are complete for request `seq` (the latest one by default). ## `handle` [Section titled “handle”](#handle) ```python handle(message: dict[str, Any]) -> None ``` Process one message from the events socket. ## `on_request` [Section titled “on_request”](#on_request) ```python on_request(callback: Callable[[], None]) -> None ``` Call `callback` (then `ack()`) whenever the node asks for a checkpoint; a no-op outside Nodus. ## `requested` [Section titled “requested”](#requested) ```python requested() -> bool ``` True while a checkpoint request is pending; poll it at safe points in a training loop. ## `subscribe` [Section titled “subscribe”](#subscribe) ```python subscribe() -> None ``` Listen for checkpoint requests without a callback: poll `requested()` and call `ack()` after saving. # nodus.cluster > `nodus.cluster.info()`: the gang env contract of distributed-training.md §3.6, as one object. `nodus.cluster.info()`: the gang env contract of distributed-training.md §3.6, as one object. It works in `@nodus.clustered` Functions and in plain gang Jobs alike; outside a gang it describes a gang of one. ## `ClusterInfo` [Section titled “ClusterInfo”](#clusterinfo) ```python class ClusterInfo(rank: int, size: int, ips: list[str], master_addr: str, master_port: int, epoch: int, transport: str | None, gpus_per_node: int) -> None ``` ### `ClusterInfo.epoch` [Section titled “ClusterInfo.epoch”](#clusterinfoepoch) Type: `int` ### `ClusterInfo.gpus_per_node` [Section titled “ClusterInfo.gpus_per_node”](#clusterinfogpus_per_node) Type: `int` ### `ClusterInfo.ips` [Section titled “ClusterInfo.ips”](#clusterinfoips) Type: `list[str]` ### `ClusterInfo.is_leader` [Section titled “ClusterInfo.is_leader”](#clusterinfois_leader) Type: `bool` Rank 0 hosts the rendezvous, and its return value is the clustered call’s result. ### `ClusterInfo.master_addr` [Section titled “ClusterInfo.master_addr”](#clusterinfomaster_addr) Type: `str` ### `ClusterInfo.master_port` [Section titled “ClusterInfo.master_port”](#clusterinfomaster_port) Type: `int` ### `ClusterInfo.rank` [Section titled “ClusterInfo.rank”](#clusterinforank) Type: `int` ### `ClusterInfo.size` [Section titled “ClusterInfo.size”](#clusterinfosize) Type: `int` ### `ClusterInfo.transport` [Section titled “ClusterInfo.transport”](#clusterinfotransport) Type: `str | None` ## `info` [Section titled “info”](#info) ```python info(env: Mapping[str, str] | None = None) -> ClusterInfo ``` # nodus.envs > Helpers for building your own Environment and datasets: `nodus.envs.split`. Helpers for building your own Environment and datasets: `nodus.envs.split`. `split` keeps train and test disjoint under a canonical identity, the rule every catalog Environment follows: two items with the same identity (the same question, the same graph up to relabeling, the same expression up to commutativity) always land in the same split, so the held-out score measures generalisation, not recall. ```plaintext train, test = nodus.envs.split(items, identity=lambda x: x["question"], test_size=0.2, seed=7) ``` ## `canonical` [Section titled “canonical”](#canonical) ```python canonical(value: Any) -> str ``` A stable identity for any JSON-like value: sha256 of its sorted, compact JSON (strings are whitespace- and case-normalised, so trivially reformatted duplicates collide). ## `split` [Section titled “split”](#split) ```python split(items: Iterable[T], identity: Callable[[T], Any] | None = None, test_size: float | int = 0.2, seed: int = 0) -> tuple[list[T], list[T]] ``` Seeded train and test lists with no identity in both; each keeps the input order. `test_size` is a fraction of the identity groups (0 < f < 1) or a number of groups. Groups, not items, are assigned, so every duplicate of a held-out item is held out too. # nodus.errors > Errors raised by the Nodus SDK. Errors raised by the Nodus SDK. Every error derives from `NodusError`. API errors map one-to-one onto the error-code registry (ADR-028): the class is chosen by the `metav1.Status` reason, so `except nodus.errors.InsufficientCredits` works the same for the CLI, the SDK and raw HTTP. Client-side failures are `APIConnectionError` and `APITimeoutError`; outcome errors are `JobFailed`, `AgentRunFailed`, `ImageBuildFailed`, `FunctionCallFailed` and `RemoteError`. ## Exports [Section titled “Exports”](#exports) * `APIConnectionError`: defined in `nodus.errors._base` * `APITimeoutError`: defined in `nodus.errors._base` * `AdmissionsPaused`: defined in `nodus.errors._codes` * `AgentRunFailed`: defined in `nodus.errors._base` * `AlreadyExists`: defined in `nodus.errors._codes` * `ArrearsOutstanding`: defined in `nodus.errors._codes` * `BY_STATUS`: defined in `nodus.errors._codes` * `BadRequest`: defined in `nodus.errors._codes` * `BudgetExceeded`: defined in `nodus.errors._codes` * `CapacityUnavailable`: defined in `nodus.errors._codes` * `Conflict`: defined in `nodus.errors._codes` * `Expired`: defined in `nodus.errors._codes` * `FieldImmutable`: defined in `nodus.errors._codes` * `Forbidden`: defined in `nodus.errors._codes` * `FunctionCallFailed`: defined in `nodus.errors._base` * `IdempotencyKeyReused`: defined in `nodus.errors._codes` * `ImageBuildFailed`: defined in `nodus.errors._base` * `InsufficientCredits`: defined in `nodus.errors._codes` * `Invalid`: defined in `nodus.errors._codes` * `JobFailed`: defined in `nodus.errors._base` * `NodusError`: defined in `nodus.errors._base` * `NotFound`: defined in `nodus.errors._codes` * `PaymentDisputed`: defined in `nodus.errors._codes` * `PaymentMethodRequired`: defined in `nodus.errors._codes` * `PaymentVerificationUnavailable`: defined in `nodus.errors._codes` * `PreconditionFailed`: defined in `nodus.errors._codes` * `QuotaExceeded`: defined in `nodus.errors._codes` * `REGISTRY`: defined in `nodus.errors._codes` * `RemoteError`: defined in `nodus.errors._base` * `RequestInProgress`: defined in `nodus.errors._codes` * `SandboxStarting`: defined in `nodus.errors._codes` * `StdinBackpressure`: defined in `nodus.errors._codes` * `TooManyRequests`: defined in `nodus.errors._codes` * `Unauthorized`: defined in `nodus.errors._codes` * `Unavailable`: defined in `nodus.errors._codes` * `Unsupported`: defined in `nodus.errors._codes` * `UnsupportedMediaType`: defined in `nodus.errors._codes` * `VolumeBusy`: defined in `nodus.errors._codes` ## `from_status` [Section titled “from_status”](#from_status) ```python from_status(http_status: int, body: Any, *, request_id: str | None = None, retry_after: float | None = None, idempotency_key: str | None = None) -> NodusError ``` Build the registry error for an HTTP error response whose body is a `metav1.Status` (or anything else). # nodus.examples > Runnable examples that ship with the SDK: `python -m nodus.examples.rl` trains a model with RL. Runnable examples that ship with the SDK: `python -m nodus.examples.rl` trains a model with RL. # nodus.examples.rl > Reinforcement learning in one file: teach a small model to spell words backwards. Reinforcement learning in one file: teach a small model to spell words backwards. Run it with `python -m nodus.examples.rl`. To train on your own task, copy this file and change `tasks` and `reward`: a task is a prompt and the answer only `reward` sees, and `reward` scores one completion against that answer. ## `main` [Section titled “main”](#main) ```python main() -> None ``` ## `reward` [Section titled “reward”](#reward) ```python reward(completion, answer) ``` 1.0 for the exact reversed word, 0.0 for a wrong one, None when the completion holds no answer. # nodus.functions > Functions, FunctionCalls and `@app.cls` classes (resources.md §3.10, §3.11, ADR-047, ADR-110). Functions, FunctionCalls and `@app.cls` classes (resources.md §3.10, §3.11, ADR-047, ADR-110). `.remote()`, `.spawn()` and `.map()` all create FunctionCall objects, the one fan-out mechanism of R3. Arguments are cloudpickled (`nodus._serialize`), results come back through a watch on the call, and a remote exception is re-raised as its own type when that type is importable here. ## `Cls` [Section titled “Cls”](#cls) ```python class Cls(fn: _Function) -> None ``` An `@app.cls` class: `Embedder()` gives an object whose `@nodus.method()`s have `.remote`, `.map`, `.spawn`. ### `Cls.from_name` [Section titled “Cls.from_name”](#clsfrom_name) ```python from_name(app: str, name: str, project: str | None = None) -> _Cls ``` ## `ClsObject` [Section titled “ClsObject”](#clsobject) ```python class ClsObject(fn: _Function) -> None ``` ## `Function` [Section titled “Function”](#function) ```python class Function(*, raw: Callable[..., Any] | None = None, app: _App | None = None, options: FunctionOptions | None = None, user_cls: type | None = None, method: str | None = None, deployed: str | None = None, deployed_app: str | None = None, project: str | None = None, instance: _Obj | None = None) -> None ``` A Function defined with `@app.function`, a method of an `@app.cls` class, or a deployed one by name. ### `Function.estimate` [Section titled “Function.estimate”](#functionestimate) ```python estimate(*args: Any, **kwargs: Any) -> View ``` Dry-run one call: cold and warm start ETAs and the rate; `.etag` binds a reviewed launch. ### `Function.for_each` [Section titled “Function.for_each”](#functionfor_each) ```python for_each(*iterables: Iterable[Any], ignore_exceptions: bool = False) -> None ``` Run the function over the inputs and discard the results. ### `Function.from_name` [Section titled “Function.from_name”](#functionfrom_name) ```python from_name(app: str, name: str, project: str | None = None) -> _Function ``` A deployed Function: `-` in the project (`Function.lookup` is an alias). ### `Function.get_raw_f` [Section titled “Function.get_raw_f”](#functionget_raw_f) ```python get_raw_f() -> Callable[..., Any] ``` The undecorated function (for a class method, the class). ### `Function.local` [Section titled “Function.local”](#functionlocal) ```python local(*args: Any, **kwargs: Any) -> Any ``` Run in this process, without Nodus. ### `Function.lookup` [Section titled “Function.lookup”](#functionlookup) ### `Function.map` [Section titled “Function.map”](#functionmap) ```python map(*iterables: Iterable[Any], kwargs: dict[str, Any] | None = None, order_outputs: bool = True, return_exceptions: bool = False) -> AsyncIterator[Any] ``` One FunctionCall per input, created in batches (≤1,000 calls, ≤8 MiB); results stream back as they finish. ### `Function.member` [Section titled “Function.member”](#functionmember) ### `Function.remote` [Section titled “Function.remote”](#functionremote) ```python remote(*args: Any, **kwargs: Any) -> Any ``` One FunctionCall: block until it finishes and return its result, or raise its exception. ### `Function.spawn` [Section titled “Function.spawn”](#functionspawn) ```python spawn(*args: Any, **kwargs: Any) -> _FunctionCall ``` Start one FunctionCall and return its handle without waiting. ### `Function.starmap` [Section titled “Function.starmap”](#functionstarmap) ```python starmap(iterable: Iterable[Iterable[Any]], order_outputs: bool = True, return_exceptions: bool = False) -> AsyncIterator[Any] ``` ## `FunctionCall` [Section titled “FunctionCall”](#functioncall) ```python class FunctionCall(name: str, function: str | None = None, project: str | None = None) -> None ``` A handle on one FunctionCall: `.get(timeout=None)`, `.cancel()`, `.status()`, `.logs()`. ### `FunctionCall.cancel` [Section titled “FunctionCall.cancel”](#functioncallcancel) ```python cancel() -> None ``` ### `FunctionCall.from_name` [Section titled “FunctionCall.from_name”](#functioncallfrom_name) ```python from_name(name: str, project: str | None = None) -> _FunctionCall ``` ### `FunctionCall.function` [Section titled “FunctionCall.function”](#functioncallfunction) Type: `str | None` ### `FunctionCall.get` [Section titled “FunctionCall.get”](#functioncallget) ```python get(timeout: float | None = None) -> Any ``` Wait for the result; raises the remote exception, or `TimeoutError` after `timeout` seconds. ### `FunctionCall.logs` [Section titled “FunctionCall.logs”](#functioncalllogs) ```python logs(follow: bool = False) -> AsyncIterator[str] ``` Log lines of the Function’s workers. ### `FunctionCall.name` [Section titled “FunctionCall.name”](#functioncallname) Type: `str` ### `FunctionCall.status` [Section titled “FunctionCall.status”](#functioncallstatus) ```python status() -> str | None ``` The call’s phase: `Queued`, `Running`, `Recovering`, `Succeeded`, `Failed`, `Cancelling` or `Cancelled`. ## `FunctionOptions` [Section titled “FunctionOptions”](#functionoptions) ```python class FunctionOptions(name: str | None = None, gpu: Any = None, cpu: Any = None, memory: Any = None, ephemeral_disk: Any = None, image: Any = None, secrets: list[Any] = list(), volumes: dict[str, Any] = dict(), env: dict[str, str] | None = None, network: Any = None, timeout: Any = None, retries: Any = None, checkpoint: Any = None, interruptible: Any = None, region: Any = None, profile: Any = None, max_cost: Any = None, min_workers: int | None = None, max_workers: int | None = None, scaledown_window: Any = None, target_concurrency: int | None = None) -> None ``` The `@app.function` / `@app.cls` arguments, mapped to Function spec fields by `spec()`. ### `FunctionOptions.checkpoint` [Section titled “FunctionOptions.checkpoint”](#functionoptionscheckpoint) Type: `Any` ### `FunctionOptions.cpu` [Section titled “FunctionOptions.cpu”](#functionoptionscpu) Type: `Any` ### `FunctionOptions.env` [Section titled “FunctionOptions.env”](#functionoptionsenv) Type: `dict[str, str] | None` ### `FunctionOptions.ephemeral_disk` [Section titled “FunctionOptions.ephemeral_disk”](#functionoptionsephemeral_disk) Type: `Any` ### `FunctionOptions.gpu` [Section titled “FunctionOptions.gpu”](#functionoptionsgpu) Type: `Any` ### `FunctionOptions.image` [Section titled “FunctionOptions.image”](#functionoptionsimage) Type: `Any` ### `FunctionOptions.interruptible` [Section titled “FunctionOptions.interruptible”](#functionoptionsinterruptible) Type: `Any` ### `FunctionOptions.max_cost` [Section titled “FunctionOptions.max_cost”](#functionoptionsmax_cost) Type: `Any` ### `FunctionOptions.max_workers` [Section titled “FunctionOptions.max_workers”](#functionoptionsmax_workers) Type: `int | None` ### `FunctionOptions.memory` [Section titled “FunctionOptions.memory”](#functionoptionsmemory) Type: `Any` ### `FunctionOptions.min_workers` [Section titled “FunctionOptions.min_workers”](#functionoptionsmin_workers) Type: `int | None` ### `FunctionOptions.name` [Section titled “FunctionOptions.name”](#functionoptionsname) Type: `str | None` ### `FunctionOptions.network` [Section titled “FunctionOptions.network”](#functionoptionsnetwork) Type: `Any` ### `FunctionOptions.profile` [Section titled “FunctionOptions.profile”](#functionoptionsprofile) Type: `Any` ### `FunctionOptions.region` [Section titled “FunctionOptions.region”](#functionoptionsregion) Type: `Any` ### `FunctionOptions.retries` [Section titled “FunctionOptions.retries”](#functionoptionsretries) Type: `Any` ### `FunctionOptions.scaledown_window` [Section titled “FunctionOptions.scaledown_window”](#functionoptionsscaledown_window) Type: `Any` ### `FunctionOptions.secrets` [Section titled “FunctionOptions.secrets”](#functionoptionssecrets) Type: `list[Any]` ### `FunctionOptions.target_concurrency` [Section titled “FunctionOptions.target_concurrency”](#functionoptionstarget_concurrency) Type: `int | None` ### `FunctionOptions.timeout` [Section titled “FunctionOptions.timeout”](#functionoptionstimeout) Type: `Any` ### `FunctionOptions.volumes` [Section titled “FunctionOptions.volumes”](#functionoptionsvolumes) Type: `dict[str, Any]` ## `clustered` [Section titled “clustered”](#clustered) ```python clustered(size: int, launcher: str = 'Plain', network: str = 'Colocated', transport: str = 'Direct') -> Callable[[Callable[..., Any]], Callable[..., Any]] ``` Run each `.remote()` or `.spawn()` as one gang of `size` nodes; rank 0’s return value is the result (Beta). Apply it below `@app.function`. `nodus.cluster.info()` gives each member its rank and the gang’s addresses. # nodus.image > The Modal-shaped `Image` builder (resources.md §5.3). The Modal-shaped `Image` builder (resources.md §5.3). A chain of immutable builder calls describes an `Image` spec. Nothing touches the network until the image is used by a published App or `.build()` is called; the Image object is named by the hash of its spec, so an identical chain reuses the built digest. ## `Image` [Section titled “Image”](#image) ```python class Image(*, base: str | None = None, dockerfile: Obj | None = None, python: str | None = None, steps: tuple[tuple[Any, ...], ...] = , existing: str | None = None, pull_secrets: tuple[str, ...] = ) -> None ``` ### `Image.add_local_dir` [Section titled “Image.add_local_dir”](#imageadd_local_dir) ```python add_local_dir(local_path: str | Path, remote_path: str) -> _Image ``` ### `Image.add_local_file` [Section titled “Image.add_local_file”](#imageadd_local_file) ```python add_local_file(local_path: str | Path, remote_path: str) -> _Image ``` ### `Image.add_local_python_source` [Section titled “Image.add_local_python_source”](#imageadd_local_python_source) ```python add_local_python_source(*modules: str) -> _Image ``` Copy importable local modules or packages to `/workspace/`, where the worker imports from. ### `Image.apt_install` [Section titled “Image.apt_install”](#imageapt_install) ```python apt_install(*packages: str) -> _Image ``` ### `Image.build` [Section titled “Image.build”](#imagebuild) ```python build(*, show_logs: bool = True) -> _Image ``` Build now (otherwise the first use builds it), streaming the build log; raises `ImageBuildFailed`. ### `Image.debian_slim` [Section titled “Image.debian_slim”](#imagedebian_slim) ```python debian_slim(python_version: str | None = None) -> _Image ``` The catalog `nodus/python:` image; defaults to the caller’s Python minor version. ### `Image.digest` [Section titled “Image.digest”](#imagedigest) Type: `str | None` The built digest, known after `.build()` or after an App that uses the image was published. ### `Image.env` [Section titled “Image.env”](#imageenv) ```python env(values: dict[str, str]) -> _Image ``` ### `Image.from_dockerfile` [Section titled “Image.from_dockerfile”](#imagefrom_dockerfile) ```python from_dockerfile(path: str | Path, context: str | Path = '.') -> _Image ``` ### `Image.from_name` [Section titled “Image.from_name”](#imagefrom_name) ```python from_name(name: str) -> _Image ``` An Image object that already exists in the project. ### `Image.from_registry` [Section titled “Image.from_registry”](#imagefrom_registry) ```python from_registry(ref: str, secret: Any = None) -> _Image ``` Start from an OCI reference or a catalog image (`nodus/pytorch:2.8-cuda12.8`). ### `Image.name` [Section titled “Image.name”](#imagename) Type: `str | None` ### `Image.pip_install` [Section titled “Image.pip_install”](#imagepip_install) ```python pip_install(*packages: str, index_url: str | None = None, extra_index_urls: list[str] | None = None) -> _Image ``` ### `Image.python` [Section titled “Image.python”](#imagepython) Type: `str | None` ### `Image.run_commands` [Section titled “Image.run_commands”](#imagerun_commands) ```python run_commands(*commands: str) -> _Image ``` ### `Image.uv_pip_install` [Section titled “Image.uv_pip_install”](#imageuv_pip_install) ```python uv_pip_install(*packages: str, index_url: str | None = None, extra_index_urls: list[str] | None = None) -> _Image ``` ### `Image.uv_sync` [Section titled “Image.uv_sync”](#imageuv_sync) ```python uv_sync(project_dir: str | Path = '.', frozen: bool = True) -> _Image ``` ### `Image.workdir` [Section titled “Image.workdir”](#imageworkdir) ```python workdir(path: str) -> _Image ``` # nodus.integrations > Framework callbacks that report training progress to Nodus from inside a container. Framework callbacks that report training progress to Nodus from inside a container. `nodus.integrations.transformers.NodusCallback` plugs into a Hugging Face `Trainer` (and every TRL trainer); `nodus.integrations.lightning.NodusCallback` into a Lightning `Trainer`. Both send metrics to `nodus.log.metrics`, report progress, save when the node asks for a checkpoint, and do nothing outside Nodus. Only the global rank 0 process reports, so a multi-node gang reports once. # nodus.integrations.lightning > `NodusCallback` for a Lightning `Trainer` (ADR-103). `NodusCallback` for a Lightning `Trainer` (ADR-103). ```python from nodus.integrations.lightning import NodusCallback trainer = L.Trainer(callbacks=[NodusCallback()], default_root_dir=nodus.state_dir()) trainer.fit(model, ckpt_path=NodusCallback.last_checkpoint()) ``` ## `NodusCallback` [Section titled “NodusCallback”](#noduscallback) ```python class NodusCallback(every_n_steps: int = 10, phase: str = 'Training') -> None ``` Metrics from `trainer.callback_metrics`, progress, phases and the checkpoint handshake. ### `NodusCallback.last_checkpoint` [Section titled “NodusCallback.last_checkpoint”](#noduscallbacklast_checkpoint) ```python last_checkpoint(directory: str | os.PathLike[str] | None = None) -> str | None ``` The checkpoint this callback saved in an earlier attempt, for `trainer.fit(ckpt_path=...)`. ### `NodusCallback.on_train_batch_end` [Section titled “NodusCallback.on_train_batch_end”](#noduscallbackon_train_batch_end) ```python on_train_batch_end(trainer: Any, pl_module: Any, outputs: Any, batch: Any, batch_idx: int) -> None ``` ### `NodusCallback.on_train_start` [Section titled “NodusCallback.on_train_start”](#noduscallbackon_train_start) ```python on_train_start(trainer: Any, pl_module: Any) -> None ``` # nodus.integrations.transformers > `NodusCallback` for the Hugging Face `Trainer` and every TRL trainer built on it (ADR-103). `NodusCallback` for the Hugging Face `Trainer` and every TRL trainer built on it (ADR-103). ```python from nodus.integrations.transformers import NodusCallback trainer = SFTTrainer(..., callbacks=[NodusCallback()]) trainer.train(resume_from_checkpoint=nodus.integrations.transformers.last_checkpoint()) ``` Point `output_dir` at `nodus.state_dir()` so the Trainer’s `checkpoint-` folders are the recovery state. ## `NodusCallback` [Section titled “NodusCallback”](#noduscallback) ```python class NodusCallback(phase: str = 'Training') -> None ``` Metrics, progress, phases and the checkpoint handshake for a `transformers.Trainer`. ### `NodusCallback.on_log` [Section titled “NodusCallback.on_log”](#noduscallbackon_log) ```python on_log(args: Any, state: Any, control: Any, logs: dict[str, Any] | None = None, **kwargs: Any) -> None ``` ### `NodusCallback.on_save` [Section titled “NodusCallback.on_save”](#noduscallbackon_save) ```python on_save(args: Any, state: Any, control: Any, **kwargs: Any) -> None ``` ### `NodusCallback.on_step_end` [Section titled “NodusCallback.on_step_end”](#noduscallbackon_step_end) ```python on_step_end(args: Any, state: Any, control: Any, **kwargs: Any) -> Any ``` ### `NodusCallback.on_train_begin` [Section titled “NodusCallback.on_train_begin”](#noduscallbackon_train_begin) ```python on_train_begin(args: Any, state: Any, control: Any, **kwargs: Any) -> None ``` ## `last_checkpoint` [Section titled “last_checkpoint”](#last_checkpoint) ```python last_checkpoint(directory: str | os.PathLike[str] | None = None) -> str | None ``` The newest complete `checkpoint-` under `directory` (default the state dir), for `resume_from_checkpoint`. A folder counts only when the Trainer finished writing it (`trainer_state.json` is present), so an attempt stopped mid-save resumes from the checkpoint before it. # nodus.job > `Job`: a run-to-completion container with checkpoints, recovery and outputs (resources.md §3.1). `Job`: a run-to-completion container with checkpoints, recovery and outputs (resources.md §3.1). `distributed=nodus.Distributed(...)` makes it a multi-node gang (Beta, ADR-112); logs and exec then take a `rank`. ## `Distributed` [Section titled “Distributed”](#distributed) ```python class Distributed(nodes: int | None = None, total_gpus: int | None = None, gpus_per_node: int | None = None, launcher: str | None = None, network: str | None = None, transport: str | None = None, startup_timeout: Any = None) -> None ``` A gang: `nodes` or `total_gpus`; `launcher` `Plain`, `Torchrun`, `Ray`, `Verl`; `network` and `transport`. ### `Distributed.gpus_per_node` [Section titled “Distributed.gpus_per_node”](#distributedgpus_per_node) Type: `int | None` ### `Distributed.launcher` [Section titled “Distributed.launcher”](#distributedlauncher) Type: `str | None` ### `Distributed.network` [Section titled “Distributed.network”](#distributednetwork) Type: `str | None` ### `Distributed.nodes` [Section titled “Distributed.nodes”](#distributednodes) Type: `int | None` ### `Distributed.spec` [Section titled “Distributed.spec”](#distributedspec) ```python spec() -> Obj ``` ### `Distributed.startup_timeout` [Section titled “Distributed.startup_timeout”](#distributedstartup_timeout) Type: `Any` ### `Distributed.total_gpus` [Section titled “Distributed.total_gpus”](#distributedtotal_gpus) Type: `int | None` ### `Distributed.transport` [Section titled “Distributed.transport”](#distributedtransport) Type: `str | None` ## `Job` [Section titled “Job”](#job) ```python class Job(obj: Obj) -> None ``` ### `Job.attempts` [Section titled “Job.attempts”](#jobattempts) ```python attempts() -> list[View] ``` ### `Job.cancel` [Section titled “Job.cancel”](#jobcancel) ```python cancel() -> None ``` ### `Job.create` [Section titled “Job.create”](#jobcreate) ```python create(*, name: str | None = None, image: Any = None, command: list[str] | None = None, args: list[str] | None = None, source: str | os.PathLike[str] | Mapping[str, Any] | None = None, gpu: Any = None, cpu: Any = None, memory: Any = None, disk: Any = None, env: dict[str, str] | None = None, secrets: list[Any] | None = None, volumes: Mapping[str, Any] | None = None, workdir: str | None = None, network: _spec.Egress | None = None, checkpoint: Any = None, timeout: Any = None, expected_duration: Any = None, interruptible: Any = None, region: Any = None, profile: Any = None, max_cost: Any = None, completions: int | None = None, parallelism: int | None = None, distributed: Distributed | None = None, outputs: Mapping[str, str] | None = None, labels: dict[str, str] | None = None, allow_large_source: bool = False, project: str | None = None) -> _Job ``` ### `Job.estimate` [Section titled “Job.estimate”](#jobestimate) ```python estimate() -> View ``` The Job’s `status.estimate`: cost p50/p90, start ETA, hold (and topology and gang hold for gangs). ### `Job.exec` [Section titled “Job.exec”](#jobexec) ```python exec(*command: str, pty: bool = False, rank: int | None = None, index: int | None = None) -> _Process ``` ### `Job.files` [Section titled “Job.files”](#jobfiles) Type: `_Files` ### `Job.from_name` [Section titled “Job.from_name”](#jobfrom_name) ```python from_name(name: str, project: str | None = None) -> _Job ``` ### `Job.logs` [Section titled “Job.logs”](#joblogs) ```python logs(follow: bool = False, rank: int | str | None = None, **params: Any) -> AsyncIterator[str] ``` Log lines; for a gang, `rank=n` reads one rank and `rank="all"` merges them with `[r]` prefixes. ### `Job.name` [Section titled “Job.name”](#jobname) Type: `str` ### `Job.outputs` [Section titled “Job.outputs”](#joboutputs) Type: `_Outputs` `job.outputs["model"].download("./model")`. ### `Job.resume` [Section titled “Job.resume”](#jobresume) ```python resume() -> None ``` ### `Job.run` [Section titled “Job.run”](#jobrun) ```python run(**kwargs: Any) -> _Job ``` Create the Job and return its handle (the same as `create`); `.wait()` blocks until it finishes. ### `Job.status` [Section titled “Job.status”](#jobstatus) ```python status() -> View ``` ### `Job.suspend` [Section titled “Job.suspend”](#jobsuspend) ```python suspend() -> None ``` ### `Job.wait` [Section titled “Job.wait”](#jobwait) ```python wait(timeout: float | None = None) -> View ``` Block until the Job finishes; raises `JobFailed` with the exit code and log tail when it fails. ## `Output` [Section titled “Output”](#output) ```python class Output(parent: Any, name: str, kind: str = 'Job') -> None ``` ### `Output.download` [Section titled “Output.download”](#outputdownload) ```python download(path: str | os.PathLike[str], index: int | None = None) -> Path ``` Save the output at `path`: the sha256 from `X-Nodus-SHA256` is verified, then the file is renamed in. ### `Output.name` [Section titled “Output.name”](#outputname) Type: `str` ## `Outputs` [Section titled “Outputs”](#outputs) ```python class Outputs(parent: Any, kind: str = 'Job') -> None ``` Declared and collected outputs of a Job, TrainingJob or AgentRun: `outputs["name"].download(path)` for one, `outputs.download(dir)` for all of them, `outputs.list()`. ### `Outputs.download` [Section titled “Outputs.download”](#outputsdownload) ```python download(path: str | os.PathLike[str], prefix: str = '') -> Path ``` Save every output named under `prefix` (all by default) into the directory `path`, each verified, at its name below `prefix`: `outputs.download("./adapter", prefix="adapter/")`. ### `Outputs.list` [Section titled “Outputs.list”](#outputslist) ```python list() -> list[Obj] ``` # nodus.llm > OpenAI- and Anthropic-compatible inference (resources.md §4.5, §4.6, ADR-049, ADR-100). OpenAI- and Anthropic-compatible inference (resources.md §4.5, §4.6, ADR-049, ADR-100). `nodus.llm.openai()` and `nodus.llm.anthropic()` return the official clients pointed at the Nodus inference data plane with your key, so every stock SDK feature works. Inside a Nodus container they use `OPENAI_BASE_URL` and `ANTHROPIC_BASE_URL`, the loopback proxy that adds the attempt’s token and bills the owning run; name an endpoint there with `model="endpoint/"`. Requires the `openai` or `anthropic` extra: `pip install "nodus-compute[openai]"`. ## `InferenceEndpoint` [Section titled “InferenceEndpoint”](#inferenceendpoint) ```python class InferenceEndpoint(obj: Obj) -> None ``` A named access policy (limits, allowed keys, cost cap) over a catalog Model, with its own base URL. ### `InferenceEndpoint.anthropic` [Section titled “InferenceEndpoint.anthropic”](#inferenceendpointanthropic) ```python anthropic(**kwargs: Any) -> Any ``` ### `InferenceEndpoint.base_url` [Section titled “InferenceEndpoint.base_url”](#inferenceendpointbase_url) Type: `str | None` ### `InferenceEndpoint.create` [Section titled “InferenceEndpoint.create”](#inferenceendpointcreate) ```python create(name: str, model: str, *, rpm: int | None = None, tpm: int | None = None, max_concurrent: int | None = None, allowed_keys: list[str] | None = None, max_cost: Any = None, project: str | None = None) -> _InferenceEndpoint ``` ### `InferenceEndpoint.from_name` [Section titled “InferenceEndpoint.from_name”](#inferenceendpointfrom_name) ```python from_name(name: str, project: str | None = None) -> _InferenceEndpoint ``` ### `InferenceEndpoint.name` [Section titled “InferenceEndpoint.name”](#inferenceendpointname) Type: `str` ### `InferenceEndpoint.openai` [Section titled “InferenceEndpoint.openai”](#inferenceendpointopenai) ```python openai(**kwargs: Any) -> Any ``` ### `InferenceEndpoint.usage` [Section titled “InferenceEndpoint.usage”](#inferenceendpointusage) ```python usage() -> View ``` Requests, errors, tokens and cost over the last 24 hours (refreshed every minute). ## `anthropic` [Section titled “anthropic”](#anthropic) ```python anthropic(endpoint: str | None = None, project: str | None = None, **kwargs: Any) -> Any ``` An `anthropic.Anthropic` client for `/v1/messages` on the inference data plane. ## `async_anthropic` [Section titled “async_anthropic”](#async_anthropic) ```python async_anthropic(endpoint: str | None = None, project: str | None = None, **kwargs: Any) -> Any ``` ## `async_openai` [Section titled “async_openai”](#async_openai) ```python async_openai(endpoint: str | None = None, project: str | None = None, **kwargs: Any) -> Any ``` ## `openai` [Section titled “openai”](#openai) ```python openai(endpoint: str | None = None, project: str | None = None, **kwargs: Any) -> Any ``` An `openai.OpenAI` client for the inference data plane (or a named InferenceEndpoint). # nodus.log > Structured telemetry from inside a container: metrics, RL task outcomes, work units and phases. Structured telemetry from inside a container: metrics, RL task outcomes, work units and phases. ## `metrics` [Section titled “metrics”](#metrics) ```python metrics(step: int | None = None, **values: float) -> None ``` Numeric metrics (`loss=0.4`); they appear in `status.progress.metrics` and the console charts. ## `phase` [Section titled “phase”](#phase) ```python phase(name: str) -> None ``` ## `task` [Section titled “task”](#task) ```python task(phase: str, task_id: str, outcome: str | None = None, reward: float | None = None, attempt: int = 1, trace: Any = None, kind: str = 'completed') -> None ``` An RL or evaluation task event; `kind="started"` marks the start. ## `task_started` [Section titled “task_started”](#task_started) ```python task_started(phase: str, task_id: str) -> None ``` ## `unit` [Section titled “unit”](#unit) ```python unit(id: str, ms: float) -> None ``` One finished unit of work and its duration, for per-unit cost and latency. # nodus.outputs > `nodus.outputs.verify()`: fail fast inside a Job when its model output would not load (resources.md §10.5). `nodus.outputs.verify()`: fail fast inside a Job when its model output would not load (resources.md §10.5). ## `verify` [Section titled “verify”](#verify) ```python verify(kind: str = 'huggingface', path: str | os.PathLike[str] | None = None) -> Path ``` Check that `path` (default `NODUS_OUTPUT_DIR`) holds a loadable `peft` adapter or `huggingface` model. # nodus.process > Processes and files inside a running Job, Sandbox or Workspace (resources.md §3.6). Processes and files inside a running Job, Sandbox or Workspace (resources.md §3.6). `stdout`, `stderr` and `stdin` use the long-poll protocol (`output`, `stdin`, `signals`), which needs no WebSocket and works through every proxy; a Process’s output is retained, so a reader that reconnects resumes from its offset. `p.attach()` opens the interactive remotecommand v5 WebSocket session instead, for terminals. ## `Attach` [Section titled “Attach”](#attach) ```python class Attach(ws: Any) -> None ``` Frames of a remotecommand v5 session: iterate for `(stream, bytes)`; `write`, `resize`, `close_stdin`. ### `Attach.close_stdin` [Section titled “Attach.close_stdin”](#attachclose_stdin) ```python close_stdin() -> None ``` ### `Attach.exit_code` [Section titled “Attach.exit_code”](#attachexit_code) Type: `int | None` Set once the session reports the process’s exit status. ### `Attach.resize` [Section titled “Attach.resize”](#attachresize) ```python resize(rows: int, cols: int) -> None ``` ### `Attach.write` [Section titled “Attach.write”](#attachwrite) ```python write(data: bytes | str) -> None ``` ## `File` [Section titled “File”](#file) ```python class File(files: _Files, path: str, mode: str) -> None ``` A file opened with `sb.open(path, mode)`: reads come from one download, writes upload on `close()`. ### `File.close` [Section titled “File.close”](#fileclose) ```python close() -> None ``` ### `File.read` [Section titled “File.read”](#fileread) ```python read(n: int = 1) -> str | bytes ``` ### `File.write` [Section titled “File.write”](#filewrite) ```python write(data: str | bytes) -> int ``` ## `Files` [Section titled “Files”](#files) ```python class Files(client: Client, kind: str, name: str, project: str | None = None) -> None ``` The `files` subresource of a Job, Sandbox or Workspace. ### `Files.list` [Section titled “Files.list”](#fileslist) ```python list(path: str = '/') -> list[Obj] ``` ### `Files.open` [Section titled “Files.open”](#filesopen) ```python open(path: str, mode: str = 'r') -> _File ``` ### `Files.read` [Section titled “Files.read”](#filesread) ```python read(path: str) -> bytes ``` ### `Files.remove` [Section titled “Files.remove”](#filesremove) ```python remove(path: str) -> None ``` ### `Files.write` [Section titled “Files.write”](#fileswrite) ```python write(path: str, data: bytes | str, expected_sha256: str | None = None) -> None ``` ## `Process` [Section titled “Process”](#process) ```python class Process(client: Client, obj: Obj) -> None ``` One command executing in a running attempt: `stdout`, `stderr`, `stdin`, `wait()`, `returncode`. ### `Process.attach` [Section titled “Process.attach”](#processattach) ```python attach(stdin: bool = True, tty: bool = False) -> AsyncIterator[_Attach] ``` An interactive session over the remotecommand v5 WebSocket (one stdin writer; many readers). ### `Process.cancel` [Section titled “Process.cancel”](#processcancel) ```python cancel() -> None ``` ### `Process.name` [Section titled “Process.name”](#processname) Type: `str` ### `Process.resize` [Section titled “Process.resize”](#processresize) ```python resize(rows: int, cols: int) -> None ``` ### `Process.returncode` [Section titled “Process.returncode”](#processreturncode) Type: `int | None` The exit code once the process has finished, else None. ### `Process.signal` [Section titled “Process.signal”](#processsignal) ```python signal(sig: str) -> None ``` Send `SIGINT`, `SIGTERM`, `SIGKILL`, `SIGHUP`, `SIGQUIT`, `SIGUSR1` or `SIGUSR2`. ### `Process.start` [Section titled “Process.start”](#processstart) ```python start(client: Client, kind: str, parent: str, command: list[str], *, tty: bool = False, stdin: bool = False, env: dict[str, str] | None = None, workdir: str | None = None, timeout: str | None = None, project: str | None = None, **params: Any) -> _Process ``` ### `Process.stderr` [Section titled “Process.stderr”](#processstderr) Type: `_StreamReader` ### `Process.stdin` [Section titled “Process.stdin”](#processstdin) Type: `_StdinWriter` ### `Process.stdout` [Section titled “Process.stdout”](#processstdout) Type: `_StreamReader` ### `Process.wait` [Section titled “Process.wait”](#processwait) ```python wait(timeout: float | None = None) -> int ``` Block until the process exits and return its exit code; `TimeoutError` after `timeout` seconds. ### `Process.write` [Section titled “Process.write”](#processwrite) ```python write(data: bytes | str) -> None ``` ## `StdinWriter` [Section titled “StdinWriter”](#stdinwriter) ```python class StdinWriter(proc: _Process) -> None ``` ### `StdinWriter.drain` [Section titled “StdinWriter.drain”](#stdinwriterdrain) ```python drain() -> None ``` Writes are sent as they are made, so there is nothing buffered to flush. ### `StdinWriter.write` [Section titled “StdinWriter.write”](#stdinwriterwrite) ```python write(data: bytes | str) -> None ``` ### `StdinWriter.write_eof` [Section titled “StdinWriter.write_eof”](#stdinwriterwrite_eof) ```python write_eof() -> None ``` ## `StreamReader` [Section titled “StreamReader”](#streamreader) ```python class StreamReader(proc: _Process, stream: str) -> None ``` stdout or stderr of a Process: `read()` returns everything; iterating yields text as it arrives. ### `StreamReader.read` [Section titled “StreamReader.read”](#streamreaderread) ```python read() -> str ``` ### `StreamReader.read_bytes` [Section titled “StreamReader.read_bytes”](#streamreaderread_bytes) ```python read_bytes() -> bytes ``` # nodus.recipes > Training recipes: TrainingJob builders for fine-tuning, pretraining, distillation, preference training and RL. Training recipes: TrainingJob builders for fine-tuning, pretraining, distillation, preference training and RL. `finetune` and `rl` build `TrainingJob`s on the catalog runtimes (resources.md §4.8, ADR-103). A builder previews with a server dry-run (`.preview()` returns a `Plan` with the estimate, the compiled Job, blocking reasons and the ETag) and runs bound to that ETag (`.run(gpu=, nodes=, max_cost=, idempotency_key=)` creates with `If-Match`). `nodes > 1` asks for a gang: the TrainingJob compiles to a Job with `spec.distributed` and the runtime launches one worker per GPU on every node with torchrun, Accelerate or DeepSpeed. ```plaintext from nodus.recipes import TrainingJob, finetune, rl job = TrainingJob.from_example("nodus/gsm8k:gsm8k-trained") print(job.preview().estimate) ``` ## Exports [Section titled “Exports”](#exports) * `Data`: defined in `nodus.recipes._job` * `LoRA`: defined in `nodus.recipes._job` * `Plan`: defined in `nodus.recipes._job` * `Run`: defined in `nodus.recipes._job` * `TrainingJob`: defined in `nodus.recipes._job` # nodus.recipes.finetune > Fine-tuning, preference training, distillation and pretraining recipes on the catalog runtimes. Fine-tuning, preference training, distillation and pretraining recipes on the catalog runtimes. Each function returns a `TrainingJob` builder; nothing runs until `.preview()` or `.run(gpu=…, max_cost=…)`. Keyword arguments not named here are runtime parameters in snake_case (`learning_rate=1e-5` is `parameters.learningRate`), validated by the server against the runtime’s schema. ```plaintext sft = finetune.sft(model="Qwen/Qwen3-1.7B", data=nodus.Volume.from_name("support-chats"), lora=finetune.LoRA(r=16), max_steps=2000).run(gpu="H100", max_cost=20) pt = finetune.pretrain(model="Qwen/Qwen3-0.6B", data=corpus, initialization="Scratch") big = finetune.sft(model=m, data=d, distributed="ZeRO3").run(gpu="H100:8", nodes=2, max_cost=200) ``` ## Exports [Section titled “Exports”](#exports) * `Data`: defined in `nodus.recipes._job` * `LoRA`: defined in `nodus.recipes._job` ## `distill` [Section titled “distill”](#distill) ```python distill(*, student: Any, teacher: Any, data: Any, temperature: float | None = None, alpha: float | None = None, lmbda: float | None = None, **params: Any) -> Any ``` Knowledge distillation from a teacher model (`distill`); `lmbda=0` is offline KD on the dataset’s text. ## `dpo` [Section titled “dpo”](#dpo) ```python dpo(*, model: Any, data: Any, beta: float | None = None, **params: Any) -> Any ``` Direct preference optimisation on prompt, chosen and rejected rows (`dpo`). ## `kto` [Section titled “kto”](#kto) ```python kto(*, model: Any, data: Any, beta: float | None = None, **params: Any) -> Any ``` KTO on prompt, completion and a boolean label (`kto`). ## `orpo` [Section titled “orpo”](#orpo) ```python orpo(*, model: Any, data: Any, beta: float | None = None, **params: Any) -> Any ``` Odds-ratio preference optimisation, no reference model (`orpo`). ## `pretrain` [Section titled “pretrain”](#pretrain) ```python pretrain(*, model: Any, data: Any, initialization: str = 'Continued', packing: bool = True, **params: Any) -> Any ``` Causal-LM pretraining on raw text (`sft`, `task: Pretrain`). `initialization="Continued"` keeps the model’s weights; `"Scratch"` uses only its config and tokenizer (optionally resized by `architecture={...}`) and starts from random weights. ## `reward` [Section titled “reward”](#reward) ```python reward(*, model: Any, data: Any, **params: Any) -> Any ``` A reward model from chosen and rejected pairs (`reward-model`). ## `sft` [Section titled “sft”](#sft) ```python sft(*, model: Any, data: Any, lora: LoRA | None = None, max_steps: int | None = None, **params: Any) -> Any ``` Supervised fine-tuning on prompt and completion (or chat `messages`) rows (`sft`). # nodus.recipes.rl > Reinforcement learning and evaluation on catalog Environments or your own. Reinforcement learning and evaluation on catalog Environments or your own. From a task list and a reward function to a training run in one call. `python -m nodus.examples.rl` runs a whole example (`nodus/examples/rl.py`), a template to copy: ```plaintext def reward(completion, answer): return 1.0 if answer in completion else 0.0 run = rl.train([("Spell 'cat' backwards.", "tac"), ...], reward, max_cost=2) # model= picks another base model run.watch() # each stage, then reward, loss and KL per step run.outputs.download("./outputs") # the LoRA adapter and the before/after comparison ``` `train` packages the reward with the statements of its file it uses as an Environment in your project, then runs GRPO on it, learning from groups of replies. `rl.train("nodus/gsm8k@1.0.0")` trains on a catalog Environment the same way. To set every GRPO parameter yourself, on a catalog Environment or one you build (`examples/training/custom-reward`): ```plaintext job = rl.grpo_lora(model="Qwen/Qwen3-0.6B", environment="nodus/graph-coloring@1.0.0", steps=50, held_out=64, seed=42) plan = job.preview() # estimate, compiled Job, blocking reasons, ETag run = plan.run(gpu="RTX-4090", max_cost=5, idempotency_key="gc-001") print(run.wait().summary.comparison) # the numbers the console shows ``` The trainer never grades itself: completions go to the platform grader and come back as task events, and the comparison is computed server-side on the same held-out tasks before and after training. ## `evaluate` [Section titled “evaluate”](#evaluate) ```python evaluate(*, model: Any, environment: str | None = None, tasks: Iterable[str] | None = None, held_out: int = 64, seed: int = 0, task_filter: Mapping[str, str] | None = None, **params: Any) -> Any ``` `mode: Evaluate` on `evaluate`: an Environment’s held-out tasks, or benchmark `tasks` such as arc_easy. ## `grpo_lora` [Section titled “grpo_lora”](#grpo_lora) ```python grpo_lora(*, model: Any, environment: str, steps: int = 50, train_tasks: int = 256, held_out: int = 64, seed: int = 0, lora: LoRA | None = None, task_filter: Mapping[str, str] | None = None, min_comparison_tasks: int = 16, **params: Any) -> Any ``` GRPO with a LoRA adapter (`grpo-lora`): baseline, train, then the final eval on the same held-out tasks. # nodus.runtime > In-container helpers that speak nodusd's sockets (resources.md §8); each is a no-op outside Nodus. In-container helpers that speak nodusd’s sockets (resources.md §8); each is a no-op outside Nodus. Telemetry goes to `/run/nodus/events.sock` as JSON lines. Without the socket, inside a Nodus container, events are printed as `nodus.event {json}` lines, which nodusd also parses; outside Nodus nothing is emitted. ## `emit` [Section titled “emit”](#emit) ```python emit(event: dict[str, Any]) -> None ``` Send one telemetry event; ids are unique per process so the node’s per-attempt dedup keeps every event. ## `in_nodus` [Section titled “in_nodus”](#in_nodus) ```python in_nodus() -> bool ``` ## `progress` [Section titled “progress”](#progress) ```python progress(completed: int | float, total: int | float | None = None) -> None ``` Report progress; it appears in `status.progress` and the console, and drives `Restartable` cursors. ## `restored` [Section titled “restored”](#restored) ```python restored() -> bool ``` True when this attempt started from a restored checkpoint. ## `self` [Section titled “self”](#self) ```python self() -> dict[str, Any] ``` `{kind, name, project, attempt, epoch, index}` of the object this container runs for, from `api.sock`. ## `state_dir` [Section titled “state_dir”](#state_dir) ```python state_dir() -> Path ``` The checkpointed directory: write model, optimizer and progress files here. # nodus.sandbox > `Sandbox`: an isolated long-running container driven by exec, files and tunnels (resources.md §3.5, ADR-044). `Sandbox`: an isolated long-running container driven by exec, files and tunnels (resources.md §3.5, ADR-044). `Sandbox.create(name=...)` is create-by-name: an identical create reconnects to (and wakes) the existing Sandbox, a different spec is `AlreadyExists` with a diff. The image defaults to `nodus/agent-tools` client-side. Spec fields the API does not serve yet (`volumes`, `ports`, `init`, `service`, egress allow-lists) raise `errors.Unsupported` before anything is sent, because the API rejects an unknown field outright. `secrets` is sent when set, and a server that does not take it yet answers `Unsupported` as well. ## `Init` [Section titled “Init”](#init) ```python class Init(git: str | None = None, ref: str | None = None, path: str | None = None, project: str | Path | None = None, setup: str | None = None) -> None ``` One-time setup: clone `git` (or upload the local `project` directory), then run `setup` as Process `setup`. ### `Init.git` [Section titled “Init.git”](#initgit) Type: `str | None` ### `Init.path` [Section titled “Init.path”](#initpath) Type: `str | None` ### `Init.project` [Section titled “Init.project”](#initproject) Type: `str | Path | None` ### `Init.ref` [Section titled “Init.ref”](#initref) Type: `str | None` ### `Init.setup` [Section titled “Init.setup”](#initsetup) Type: `str | None` ## `Sandbox` [Section titled “Sandbox”](#sandbox) ```python class Sandbox(obj: Obj) -> None ``` ### `Sandbox.create` [Section titled “Sandbox.create”](#sandboxcreate) ```python create(*command: str, name: str | None = None, image: Any = None, cpu: Any = None, memory: Any = None, gpu: Any = None, disk: Any = None, timeout: Any = None, idle_timeout: Any = None, on_idle: str = 'stop', secrets: list[Any] | None = None, volumes: Mapping[str, Any] | None = None, env: dict[str, str] | None = None, workdir: str | None = None, network: _spec.Egress | None = None, ports: list[int] | None = None, init: Init | None = None, service: Service | None = None, max_cost: Any = None, labels: dict[str, str] | None = None, idempotency_key: str | None = None, project: str | None = None) -> _Sandbox ``` Create (or reconnect to, by `name`) a Sandbox; returns without waiting for it to start. ### `Sandbox.exec` [Section titled “Sandbox.exec”](#sandboxexec) ```python exec(*command: str, pty: bool = False, stdin: bool = False, env: dict[str, str] | None = None, workdir: str | None = None, timeout: Any = None) -> _Process ``` Start a recorded Process; a single string with spaces runs under `/bin/sh -c`. Wakes a stopped Sandbox. ### `Sandbox.files` [Section titled “Sandbox.files”](#sandboxfiles) Type: `_Files` ### `Sandbox.from_id` [Section titled “Sandbox.from_id”](#sandboxfrom_id) ```python from_id(object_id: str, project: str | None = None) -> _Sandbox ``` ### `Sandbox.from_name` [Section titled “Sandbox.from_name”](#sandboxfrom_name) ```python from_name(name: str, project: str | None = None) -> _Sandbox ``` Reconnect by name; a Sandbox the user stopped is started again. ### `Sandbox.list` [Section titled “Sandbox.list”](#sandboxlist) ```python list(labels: Mapping[str, str] | None = None, project: str | None = None) -> list[_Sandbox] ``` ### `Sandbox.logs` [Section titled “Sandbox.logs”](#sandboxlogs) ```python logs(follow: bool = False) -> AsyncIterator[str] ``` ### `Sandbox.name` [Section titled “Sandbox.name”](#sandboxname) Type: `str` ### `Sandbox.object_id` [Section titled “Sandbox.object_id”](#sandboxobject_id) Type: `str | None` ### `Sandbox.open` [Section titled “Sandbox.open”](#sandboxopen) ```python open(path: str, mode: str = 'r') -> _File ``` ### `Sandbox.refresh_secrets` [Section titled “Sandbox.refresh_secrets”](#sandboxrefresh_secrets) ```python refresh_secrets() -> None ``` ### `Sandbox.snapshot_filesystem` [Section titled “Sandbox.snapshot_filesystem”](#sandboxsnapshot_filesystem) ```python snapshot_filesystem(name: str | None = None) -> Any ``` Build an Image from this Sandbox’s filesystem (Beta); use it as `image=` for new Sandboxes. ### `Sandbox.start` [Section titled “Sandbox.start”](#sandboxstart) ```python start() -> None ``` ### `Sandbox.status` [Section titled “Sandbox.status”](#sandboxstatus) ```python status() -> View ``` The Sandbox’s `status` (phase, activity, endpoints, stop reason, cost). ### `Sandbox.stop` [Section titled “Sandbox.stop”](#sandboxstop) ```python stop() -> None ``` ### `Sandbox.terminate` [Section titled “Sandbox.terminate”](#sandboxterminate) ```python terminate() -> None ``` ### `Sandbox.tunnels` [Section titled “Sandbox.tunnels”](#sandboxtunnels) Type: `_Tunnels` ### `Sandbox.wait` [Section titled “Sandbox.wait”](#sandboxwait) ```python wait() -> int | None ``` Wait for the main command to exit and return its exit code (None when the Sandbox stopped without one). ## `Service` [Section titled “Service”](#service) ```python class Service(command: str | list[str], port: int, health_path: str | None = None) -> None ``` A supervised main service: restarted on exit, health-checked on `health_path`. ### `Service.command` [Section titled “Service.command”](#servicecommand) Type: `str | list[str]` ### `Service.health_path` [Section titled “Service.health_path”](#servicehealth_path) Type: `str | None` ### `Service.port` [Section titled “Service.port”](#serviceport) Type: `int` ## `Tunnel` [Section titled “Tunnel”](#tunnel) ```python class Tunnel(port: int, url: str, public: bool = False) -> None ``` ### `Tunnel.port` [Section titled “Tunnel.port”](#tunnelport) Type: `int` ### `Tunnel.public` [Section titled “Tunnel.public”](#tunnelpublic) Type: `bool` ### `Tunnel.url` [Section titled “Tunnel.url”](#tunnelurl) Type: `str` ## `Tunnels` [Section titled “Tunnels”](#tunnels) ```python class Tunnels(sb: _Sandbox) -> None ``` `sb.tunnels.open(port)` adds an ingress port and returns its URL; `sb.tunnels()` lists them. ### `Tunnels.open` [Section titled “Tunnels.open”](#tunnelsopen) ```python open(port: int, public: bool = False) -> Tunnel ``` # nodus.secret > `Secret` references (resources.md §5.2). Values are write-only: nothing the API returns contains them. `Secret` references (resources.md §5.2). Values are write-only: nothing the API returns contains them. ## `Secret` [Section titled “Secret”](#secret) ```python class Secret(name: str | None = None, *, data: dict[str, str] | None = None, type: str = 'Opaque') -> None ``` A named Secret, or inline values that become one when first used. `from_name` never touches the network. `from_dict` and `from_dotenv` are stored as `-` each time an App using them is published (owned by that App object and removed with it), or as `sdk-` outside an App. ### `Secret.create` [Section titled “Secret.create”](#secretcreate) ```python create(name: str, data: dict[str, str], type: str = 'Opaque', project: str | None = None) -> _Secret ``` Create the Secret, or write a new version when it exists (running attempts keep their pinned version). ### `Secret.from_dict` [Section titled “Secret.from_dict”](#secretfrom_dict) ```python from_dict(data: dict[str, str]) -> _Secret ``` ### `Secret.from_dotenv` [Section titled “Secret.from_dotenv”](#secretfrom_dotenv) ```python from_dotenv(path: str | Path = '.env') -> _Secret ``` ### `Secret.from_name` [Section titled “Secret.from_name”](#secretfrom_name) ```python from_name(name: str) -> _Secret ``` ### `Secret.name` [Section titled “Secret.name”](#secretname) Type: `str | None` # nodus.sweep > `Sweep`: one Job template run across a matrix of GPUs, regions and parameters (resources.md §3.3, ADR-051). `Sweep`: one Job template run across a matrix of GPUs, regions and parameters (resources.md §3.3, ADR-051). `Sweep(job_spec, grid={"gpu": ["H100", "L4"], "BATCH_SIZE": [8, 16]})` describes the matrix and `run()` creates it; each cell is a Job named `-` and the report lists cost, wall time and throughput per cell. The API serves Job templates, so a Function or a recipe as the target raises `errors.Unsupported`. ## `Sweep` [Section titled “Sweep”](#sweep) ```python class Sweep(target: Mapping[str, Any] | None = None, *, grid: Mapping[str, Iterable[Any] | str] | None = None, repetitions: int | None = None, max_parallel: int | None = None, max_cost: Any = None, name: str | None = None, labels: dict[str, str] | None = None, project: str | None = None) -> None ``` A Sweep: describe it, `run()` it, then read the report with `wait()`, `cells()` and `status()`. ### `Sweep.cancel` [Section titled “Sweep.cancel”](#sweepcancel) ```python cancel() -> None ``` ### `Sweep.cells` [Section titled “Sweep.cells”](#sweepcells) ```python cells() -> list[View] ``` One entry per cell, in index order: GPU, region, parameters, phase, cost, wall time and throughput. ### `Sweep.estimate` [Section titled “Sweep.estimate”](#sweepestimate) ```python estimate() -> View ``` The Sweep’s `status.estimate`: the total cost estimate over every cell. ### `Sweep.from_name` [Section titled “Sweep.from_name”](#sweepfrom_name) ```python from_name(name: str, project: str | None = None) -> _Sweep ``` A Sweep that exists already, to wait on or read. ### `Sweep.jobs` [Section titled “Sweep.jobs”](#sweepjobs) ```python jobs() -> list[View] ``` The cell Jobs the Sweep created. ### `Sweep.name` [Section titled “Sweep.name”](#sweepname) Type: `str | None` The Sweep’s name; a Sweep without one gets it from `run()`. ### `Sweep.resume` [Section titled “Sweep.resume”](#sweepresume) ```python resume() -> None ``` ### `Sweep.run` [Section titled “Sweep.run”](#sweeprun) ```python run(idempotency_key: str | None = None) -> _Sweep ``` Create the Sweep and return it; `.wait()` blocks until every cell is terminal. ### `Sweep.status` [Section titled “Sweep.status”](#sweepstatus) ```python status() -> View ``` The Sweep’s `status`: phase, cells, best, cost and estimate. ### `Sweep.suspend` [Section titled “Sweep.suspend”](#sweepsuspend) ```python suspend() -> None ``` ### `Sweep.wait` [Section titled “Sweep.wait”](#sweepwait) ```python wait(timeout: float | None = None) -> View ``` Block until the Sweep is terminal and return its report (the `status`). A Sweep with failed cells ends `Failed` with reason `CellsFailed` and still has its report, so only a timeout or a missing Sweep raises; read `phase` and `best` on the result. # nodus.volume > `Volume`: named storage with Modal's commit and reload semantics for `ReadWriteMany` (resources.md §5.1, ADR-091). `Volume`: named storage with Modal’s commit and reload semantics for `ReadWriteMany` (resources.md §5.1, ADR-091). File transfers go through the kopia client inside the `nodus` CLI with a storage grant from the Volume’s `uploads` subresource, so they run at object-store speed and never proxy bytes through the API. ## `BatchUpload` [Section titled “BatchUpload”](#batchupload) ```python class BatchUpload(volume: _Volume) -> None ``` ### `BatchUpload.put_directory` [Section titled “BatchUpload.put_directory”](#batchuploadput_directory) ```python put_directory(local_path: str | os.PathLike[str], remote_path: str) -> None ``` ### `BatchUpload.put_file` [Section titled “BatchUpload.put_file”](#batchuploadput_file) ```python put_file(local_path: str | os.PathLike[str], remote_path: str) -> None ``` ## `Volume` [Section titled “Volume”](#volume) ```python class Volume(name: str, *, create_if_missing: bool = False, access_mode: str | None = None, size: str = '50Gi', project: str | None = None, source: Obj | None = None) -> None ``` ### `Volume.batch_upload` [Section titled “Volume.batch_upload”](#volumebatch_upload) ```python batch_upload() -> _BatchUpload ``` `with vol.batch_upload() as up: up.put_file(...); up.put_directory(...)` uploads as one revision. ### `Volume.clear` [Section titled “Volume.clear”](#volumeclear) ```python clear() -> None ``` Commit an empty revision; the Volume and its history stay. ### `Volume.commit` [Section titled “Volume.commit”](#volumecommit) ```python commit() -> None ``` Inside a container that mounts this Volume: publish this attempt’s changed paths as a new revision. ### `Volume.from_name` [Section titled “Volume.from_name”](#volumefrom_name) ```python from_name(name: str, create_if_missing: bool = False, access_mode: str | None = None, size: str = '50Gi', project: str | None = None) -> _Volume ``` A reference resolved on first use; `create_if_missing=True` creates a `ReadWriteMany` Volume (as Modal). ### `Volume.import_from` [Section titled “Volume.import_from”](#volumeimport_from) ```python import_from(name: str, *, huggingface: str | None = None, revision: str | None = None, git: str | None = None, url: str | None = None, sha256: str | None = None, extract: str = 'auto', s3: str | None = None, connection: str | None = None, query: str | None = None, secret: str | None = None, size: str = '50Gi', project: str | None = None) -> _Volume ``` Create a `ReadOnlyMany` Volume imported from Hugging Face, git, a URL, S3 or a Connection query. ### `Volume.listdir` [Section titled “Volume.listdir”](#volumelistdir) ```python listdir(path: str = '/') -> list[str] ``` ### `Volume.name` [Section titled “Volume.name”](#volumename) Type: `str` ### `Volume.put_directory` [Section titled “Volume.put_directory”](#volumeput_directory) ```python put_directory(local_path: str | os.PathLike[str], remote_path: str) -> None ``` ### `Volume.put_file` [Section titled “Volume.put_file”](#volumeput_file) ```python put_file(local_path: str | os.PathLike[str], remote_path: str) -> None ``` ### `Volume.read_file` [Section titled “Volume.read_file”](#volumeread_file) ```python read_file(remote_path: str) -> bytes ``` ### `Volume.reload` [Section titled “Volume.reload”](#volumereload) ```python reload() -> None ``` Inside a container: remount the latest revision (close open files first). ### `Volume.remove_file` [Section titled “Volume.remove_file”](#volumeremove_file) ```python remove_file(path: str) -> None ``` ### `Volume.revisions` [Section titled “Volume.revisions”](#volumerevisions) ```python revisions() -> list[Obj] ``` # nodus.workspace > `Workspace`: a development machine with SSH, VS Code, JupyterLab and a persistent home (resources.md §3.8). `Workspace`: a development machine with SSH, VS Code, JupyterLab and a persistent home (resources.md §3.8). ## `Workspace` [Section titled “Workspace”](#workspace) ```python class Workspace(obj: Obj) -> None ``` ### `Workspace.create` [Section titled “Workspace.create”](#workspacecreate) ```python create(name: str, *, gpu: Any = None, cpu: Any = None, memory: Any = None, disk: Any = None, image: str | None = None, volume: str | _Volume | None = None, ephemeral: bool = False, tools: list[str] | None = None, idle_timeout: Any = None, secrets: list[Any] | None = None, env: dict[str, str] | None = None, max_cost: Any = None, project: str | None = None) -> _Workspace ``` Create a Workspace; unless `ephemeral=True` its home is `volume` (default `-home`), made if missing. ### `Workspace.exec` [Section titled “Workspace.exec”](#workspaceexec) ```python exec(*command: str, pty: bool = False) -> _Process ``` ### `Workspace.files` [Section titled “Workspace.files”](#workspacefiles) Type: `_Files` ### `Workspace.from_name` [Section titled “Workspace.from_name”](#workspacefrom_name) ```python from_name(name: str, project: str | None = None) -> _Workspace ``` ### `Workspace.logs` [Section titled “Workspace.logs”](#workspacelogs) ```python logs(follow: bool = False) -> AsyncIterator[str] ``` ### `Workspace.name` [Section titled “Workspace.name”](#workspacename) Type: `str` ### `Workspace.open` [Section titled “Workspace.open”](#workspaceopen) ```python open(tool: str = 'vscode', browser: bool = True) -> str ``` Mint a preview URL for VS Code or JupyterLab in the browser, open it, and return it. ### `Workspace.schedule` [Section titled “Workspace.schedule”](#workspaceschedule) ```python schedule(ready_by: str | datetime, stop_at: str | datetime | None = None) -> None ``` Be ready by `ready_by` (it starts 15 minutes earlier) and stop at `stop_at`. ### `Workspace.sessions` [Section titled “Workspace.sessions”](#workspacesessions) ```python sessions() -> list[View] ``` Every running period of this Workspace, each an Attempt whose status is the session’s receipt. ### `Workspace.ssh` [Section titled “Workspace.ssh”](#workspacessh) ```python ssh() -> None ``` Open an interactive SSH session through the CLI’s `ssh-proxy` (the CLI writes `~/.ssh/config`). ### `Workspace.start` [Section titled “Workspace.start”](#workspacestart) ```python start() -> None ``` ### `Workspace.status` [Section titled “Workspace.status”](#workspacestatus) ```python status() -> View ``` ### `Workspace.stop` [Section titled “Workspace.stop”](#workspacestop) ```python stop() -> None ``` ### `Workspace.unschedule` [Section titled “Workspace.unschedule”](#workspaceunschedule) ```python unschedule() -> None ``` ### `Workspace.wait_ready` [Section titled “Workspace.wait_ready”](#workspacewait_ready) ```python wait_ready(timeout: float | None = None) -> View ``` Block until the tools answer (phase `Running`); raises when it fails or stops instead. # Training runtimes > Every catalog TrainingRuntime, its task and modes. Beta TrainingRuntimes are Beta. The catalog runtimes in project `nodus`; reference one as `spec.runtime: nodus/` on a TrainingJob. | Runtime | Task | Modes | Summary | | -------------------------------------------------------- | -------- | --------------- | ------------------------------------------------------------------------------------------------ | | [`distill`](/docs/reference/runtimes/distill/) | Distill | Train | Knowledge distillation from a teacher model; lmbda 0 is offline distillation | | [`dpo`](/docs/reference/runtimes/dpo/) | DPO | Train | Direct preference optimization on prompt, chosen and rejected rows | | [`evaluate`](/docs/reference/runtimes/evaluate/) | Evaluate | Evaluate | Benchmark evaluation on standard tasks, or over an Environment’s held-out tasks | | [`grpo-lora`](/docs/reference/runtimes/grpo-lora/) | GRPO | Train, Evaluate | GRPO reinforcement learning with a LoRA adapter over an Environment, graded by the platform | | [`kto`](/docs/reference/runtimes/kto/) | KTO | Train | Kahneman-Tversky optimization on unpaired desirable and undesirable completions | | [`orpo`](/docs/reference/runtimes/orpo/) | ORPO | Train | Odds-ratio preference optimization: SFT and preference alignment in one pass, no reference model | | [`reward-model`](/docs/reference/runtimes/reward-model/) | Reward | Train | Reward-model training: a one-logit head scored on chosen versus rejected | | [`sft`](/docs/reference/runtimes/sft/) | SFT | Train | Supervised fine-tuning (full, LoRA, QLoRA) and pretraining with task: Pretrain | # distill > Knowledge distillation from a teacher model; lmbda 0 is offline distillation Beta TrainingJob and TrainingRuntime are Beta: fields can still change before GA. Knowledge distillation from a teacher model; lmbda 0 is offline distillation. | Field | Value | | ----------- | ----------------------------------- | | Reference | `nodus/distill` | | Image | `nodus/distill:2.0.0` | | Task | Distill | | Modes | Train | | Launcher | Torchrun | | Requires | model yes, data yes, environment no | | Checkpoints | `/nodus/state` (HFTrainer) | ## Parameters [Section titled “Parameters”](#parameters) Set under `spec.parameters`; the server validates them against this schema, and unset parameters take the runtime default. | Parameter | Type | Default | Allowed | Description | | ---------------------------------- | ------- | --------- | ------------------------------------------------------ | ------------------------------------------------------------------------------------------- | | `alpha` | number | `0.5` | 0 to 1 | Generalized JSD interpolation: 0 forward KL, 1 reverse KL | | `attention` | string | none | `eager`, `sdpa`, `flash_attention_2` | Attention kernel | | `batchSize` | integer | `4` | 1 to 1024 | Examples per GPU per step | | `distributed.findUnusedParameters` | boolean | none | | DDP: tolerate parameters without gradients | | `distributed.strategy` | string | `Auto` | `Auto`, `DDP`, `FSDP`, `ZeRO1`, `ZeRO2`, `ZeRO3` | Auto and DDP replicate the model; FSDP and ZeRO shard it | | `epochs` | number | `1` | 0.01 to 100 | Passes over the data when `steps` is -1 | | `evalSteps` | integer | none | 1 to 100000 | Steps between validation passes; unset evaluates at the end only | | `gradientAccumulation` | integer | `4` | 1 to 1024 | Steps whose gradients are summed before an update | | `gradientCheckpointing` | boolean | `true` | | Recompute activations to save GPU memory | | `learningRate` | number | `0.00001` | 0 to 1 | Peak learning rate | | `lmbda` | number | `0` | 0 to 1 | Fraction of batches whose completions the student samples itself; 0 is offline distillation | | `loggingSteps` | integer | `10` | 1 to 10000 | Steps between metric reports | | `lora.alpha` | integer | `32` | 1 to 4096 | Adapter scaling numerator | | `lora.dropout` | number | `0.05` | 0 to 0.9 | Adapter dropout | | `lora.r` | integer | `16` | 1 to 1024 | Adapter rank | | `lora.targetModules` | any | none | | Module names, or all-linear | | `lrScheduler` | string | `cosine` | `cosine`, `linear`, `constant`, `constant_with_warmup` | Learning-rate schedule | | `maxGradNorm` | number | none | 0 to 100 | Gradient clipping norm | | `maxLength` | integer | `1024` | 16 to 131072 | Tokens per example after truncation (packed block size for pretraining) | | `maxNewTokens` | integer | `128` | 1 to 8192 | Most new tokens per student-sampled completion | | `method` | string | `LoRA` | `Full`, `LoRA`, `QLoRA` | Full fine-tuning, a LoRA adapter, or LoRA on a 4-bit quantized base (GPU only) | | `precision` | string | `Auto` | `Auto`, `bf16`, `fp16`, `fp32` | Numeric precision; Auto is bf16 where the GPU supports it, else fp16 | | `saveSteps` | integer | `200` | 1 to 100000 | Steps between recovery checkpoints | | `saveTotalLimit` | integer | none | 1 to 10 | Recovery checkpoints kept in the state directory | | `seed` | integer | `42` | 0 to 2147483647 | Seed for data order, splits and initialization | | `sequenceKD` | boolean | none | | Train on teacher-generated sequences (sequence-level KD) | | `steps` | integer | `-1` | -1 to 1000000 | Optimizer steps; -1 trains for `epochs` instead | | `temperature` | number | `2` | 0.05 to 20 | Softmax temperature of teacher and student | | `warmupRatio` | number | `0.03` | 0 to 0.5 | Fraction of steps spent warming up the learning rate | | `weightDecay` | number | none | 0 to 1 | AdamW weight decay | ## Outputs [Section titled “Outputs”](#outputs) | Output | Path | | --------- | ---------------- | | `outputs` | `/nodus/outputs` | Every run also writes `results.json`, `provenance.json` and a sha256 `manifest.json`. Each file under `/nodus/outputs` is an output of the TrainingJob, named by its path: download one with `nodus cp tj/:outputs/results.json ./results.json`. ## Presets [Section titled “Presets”](#presets) Without `spec.resources`, the TrainingJob takes the resources of the preset matching its model, method and quantization, else the runtime default. | Model | Method | Quantization | GPUs | | --------------- | ------ | ------------ | --------------------------------- | | Qwen/Qwen3-0.6B | LoRA | none | 1 × RTX-4090 or L40S or RTX-A6000 | | Qwen/Qwen3-0.6B | Full | none | 1 × L40S or RTX-A6000 or A100-80G | | Qwen/Qwen3-1.7B | LoRA | none | 1 × RTX-4090 or L40S or RTX-A6000 | | Qwen/Qwen3-1.7B | Full | none | 1 × A100-80G or H100 | | Qwen/Qwen3-4B | QLoRA | nf4 | 1 × RTX-4090 or L40S or RTX-A6000 | | Qwen/Qwen3-4B | LoRA | none | 1 × L40S or RTX-A6000 or A100-80G | | Qwen/Qwen3-8B | QLoRA | nf4 | 1 × L40S or RTX-A6000 or A100-80G | | Qwen/Qwen3-8B | Full | none | 8 × H100 | Default resources: 1 × RTX-4090 or L40S or RTX-A6000 or A100-80G. ## Estimates [Section titled “Estimates”](#estimates) The dry-run estimate multiplies these by the run’s steps or tasks; the nightly runs re-measure them. | Accelerator | Seconds | Per | | ----------- | ------- | -------------- | | A100-80G | 0.7 | secondsPerStep | | H100 | 0.45 | secondsPerStep | | L40S | 1 | secondsPerStep | | RTX-4090 | 1.2 | secondsPerStep | | RTX-A6000 | 1.3 | secondsPerStep | ## Example [Section titled “Example”](#example) ```yaml apiVersion: nodus.dev/v1beta1 kind: TrainingJob metadata: {name: my-distill} spec: runtime: nodus/distill model: {uri: "hf://Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca"} data: {volume: my-dataset, format: JSONL} maxCostUSD: "5.00" ``` # dpo > Direct preference optimization on prompt, chosen and rejected rows Beta TrainingJob and TrainingRuntime are Beta: fields can still change before GA. Direct preference optimization on prompt, chosen and rejected rows. | Field | Value | | ----------- | ----------------------------------- | | Reference | `nodus/dpo` | | Image | `nodus/dpo:2.0.0` | | Task | DPO | | Modes | Train | | Launcher | Torchrun | | Requires | model yes, data yes, environment no | | Checkpoints | `/nodus/state` (HFTrainer) | ## Parameters [Section titled “Parameters”](#parameters) Set under `spec.parameters`; the server validates them against this schema, and unset parameters take the runtime default. | Parameter | Type | Default | Allowed | Description | | ---------------------------------- | ------- | ---------- | ------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------ | | `attention` | string | none | `eager`, `sdpa`, `flash_attention_2` | Attention kernel | | `batchSize` | integer | `4` | 1 to 1024 | Examples per GPU per step | | `beta` | number | `0.1` | 0 to 10 | Strength of the KL penalty toward the reference model | | `distributed.findUnusedParameters` | boolean | none | | DDP: tolerate parameters without gradients | | `distributed.strategy` | string | `Auto` | `Auto`, `DDP`, `FSDP`, `ZeRO1`, `ZeRO2`, `ZeRO3` | Auto and DDP replicate the model; FSDP and ZeRO shard it | | `epochs` | number | `1` | 0.01 to 100 | Passes over the data when `steps` is -1 | | `evalSteps` | integer | none | 1 to 100000 | Steps between validation passes; unset evaluates at the end only | | `gradientAccumulation` | integer | `4` | 1 to 1024 | Steps whose gradients are summed before an update | | `gradientCheckpointing` | boolean | `true` | | Recompute activations to save GPU memory | | `learningRate` | number | `0.000005` | 0 to 1 | Peak learning rate | | `loggingSteps` | integer | `10` | 1 to 10000 | Steps between metric reports | | `lora.alpha` | integer | `32` | 1 to 4096 | Adapter scaling numerator | | `lora.dropout` | number | `0.05` | 0 to 0.9 | Adapter dropout | | `lora.r` | integer | `16` | 1 to 1024 | Adapter rank | | `lora.targetModules` | any | none | | Module names, or all-linear | | `lossType` | string | `sigmoid` | `sigmoid`, `ipo`, `hinge`, `robust`, `sppo_hard`, `nca_pair`, `apo_zero`, `apo_down` | DPO loss | | `lrScheduler` | string | `cosine` | `cosine`, `linear`, `constant`, `constant_with_warmup` | Learning-rate schedule | | `maxGradNorm` | number | none | 0 to 100 | Gradient clipping norm | | `maxLength` | integer | `1024` | 16 to 131072 | Tokens per example after truncation (packed block size for pretraining) | | `method` | string | `LoRA` | `Full`, `LoRA`, `QLoRA` | Full fine-tuning, a LoRA adapter, or LoRA on a 4-bit quantized base (GPU only) | | `precision` | string | `Auto` | `Auto`, `bf16`, `fp16`, `fp32` | Numeric precision; Auto is bf16 where the GPU supports it, else fp16 | | `saveSteps` | integer | `200` | 1 to 100000 | Steps between recovery checkpoints | | `saveTotalLimit` | integer | none | 1 to 10 | Recovery checkpoints kept in the state directory | | `seed` | integer | `42` | 0 to 2147483647 | Seed for data order, splits and initialization | | `steps` | integer | `-1` | -1 to 1000000 | Optimizer steps; -1 trains for `epochs` instead | | `warmupRatio` | number | `0.03` | 0 to 0.5 | Fraction of steps spent warming up the learning rate | | `weightDecay` | number | none | 0 to 1 | AdamW weight decay | ## Outputs [Section titled “Outputs”](#outputs) | Output | Path | | --------- | ---------------- | | `outputs` | `/nodus/outputs` | Every run also writes `results.json`, `provenance.json` and a sha256 `manifest.json`. Each file under `/nodus/outputs` is an output of the TrainingJob, named by its path: download one with `nodus cp tj/:outputs/results.json ./results.json`. ## Presets [Section titled “Presets”](#presets) Without `spec.resources`, the TrainingJob takes the resources of the preset matching its model, method and quantization, else the runtime default. | Model | Method | Quantization | GPUs | | --------------- | ------ | ------------ | --------------------------------- | | Qwen/Qwen3-0.6B | LoRA | none | 1 × RTX-4090 or L40S or RTX-A6000 | | Qwen/Qwen3-0.6B | Full | none | 1 × L40S or RTX-A6000 or A100-80G | | Qwen/Qwen3-1.7B | LoRA | none | 1 × RTX-4090 or L40S or RTX-A6000 | | Qwen/Qwen3-1.7B | Full | none | 1 × A100-80G or H100 | | Qwen/Qwen3-4B | QLoRA | nf4 | 1 × RTX-4090 or L40S or RTX-A6000 | | Qwen/Qwen3-4B | LoRA | none | 1 × L40S or RTX-A6000 or A100-80G | | Qwen/Qwen3-8B | QLoRA | nf4 | 1 × L40S or RTX-A6000 or A100-80G | | Qwen/Qwen3-8B | Full | none | 8 × H100 | Default resources: 1 × RTX-4090 or L40S or RTX-A6000 or A100-80G. ## Estimates [Section titled “Estimates”](#estimates) The dry-run estimate multiplies these by the run’s steps or tasks; the nightly runs re-measure them. | Accelerator | Seconds | Per | | ----------- | ------- | -------------- | | A100-80G | 0.7 | secondsPerStep | | H100 | 0.45 | secondsPerStep | | L40S | 1 | secondsPerStep | | RTX-4090 | 1.2 | secondsPerStep | | RTX-A6000 | 1.3 | secondsPerStep | ## Example [Section titled “Example”](#example) ```yaml apiVersion: nodus.dev/v1beta1 kind: TrainingJob metadata: {name: my-dpo} spec: runtime: nodus/dpo model: {uri: "hf://Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca"} data: {volume: my-dataset, format: JSONL} maxCostUSD: "5.00" ``` # evaluate > Benchmark evaluation on standard tasks, or over an Environment's held-out tasks Beta TrainingJob and TrainingRuntime are Beta: fields can still change before GA. Benchmark evaluation on standard tasks, or over an Environment’s held-out tasks. | Field | Value | | ----------- | ---------------------------------- | | Reference | `nodus/evaluate` | | Image | `nodus/evaluate:2.0.0` | | Task | Evaluate | | Modes | Evaluate | | Launcher | Torchrun | | Requires | model yes, data no, environment no | | Checkpoints | `/nodus/state` (HFTrainer) | ## Parameters [Section titled “Parameters”](#parameters) Set under `spec.parameters`; the server validates them against this schema, and unset parameters take the runtime default. | Parameter | Type | Default | Allowed | Description | | --------------------- | --------------- | ------- | ---------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `batchSize` | integer | `8` | 1 to 1024 | Examples per forward pass | | `evaluationBatchSize` | integer | none | 1 to 1024 | Prompts generated together during Environment baseline and final evaluation | | `limit` | integer | none | 1 to 1000000 | Samples per task (unset scores the whole task) | | `maxCompletionLength` | integer | none | 8 to 32768 | Environment evaluation: tokens per completion | | `numFewshot` | integer | none | 0 to 64 | Few-shot examples per prompt | | `seed` | integer | `0` | 0 to 2147483647 | Seed for sampling and few-shot selection | | `taskFilter` | object | none | at most 8 keys | Keep only tasks whose metadata has these values (such as family on reasoning-gym), before trainTasks and heldOutTasks count them | | `tasks` | array of string | none | at most 64 items | Benchmark task names, such as arc_easy or hellaswag, whose datasets the run downloads from the Hugging Face Hub before it scores them offline; set either tasks or spec.environment, which scores an Environment’s held-out tasks instead | ## Outputs [Section titled “Outputs”](#outputs) | Output | Path | | --------- | ---------------- | | `outputs` | `/nodus/outputs` | Every run also writes `results.json`, `provenance.json` and a sha256 `manifest.json`. Each file under `/nodus/outputs` is an output of the TrainingJob, named by its path: download one with `nodus cp tj/:outputs/results.json ./results.json`. Default resources: 1 × RTX-4090 or L40S or RTX-A6000 or A100-80G. ## Estimates [Section titled “Estimates”](#estimates) The dry-run estimate multiplies these by the run’s steps or tasks; the nightly runs re-measure them. | Accelerator | Seconds | Per | | ----------- | ------- | -------------- | | A100-80G | 0.17 | secondsPerTask | | H100 | 0.11 | secondsPerTask | | L40S | 0.25 | secondsPerTask | | RTX-4090 | 0.3 | secondsPerTask | | RTX-A6000 | 0.33 | secondsPerTask | ## Example [Section titled “Example”](#example) ```yaml apiVersion: nodus.dev/v1beta1 kind: TrainingJob metadata: {name: my-evaluate} spec: runtime: nodus/evaluate mode: Evaluate model: {uri: "hf://Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca"} parameters: {tasks: [arc_easy], limit: 200} maxCostUSD: "5.00" ``` # grpo-lora > GRPO reinforcement learning with a LoRA adapter over an Environment, graded by the platform Beta TrainingJob and TrainingRuntime are Beta: fields can still change before GA. GRPO reinforcement learning with a LoRA adapter over an Environment, graded by the platform. | Field | Value | | ----------- | ----------------------------------- | | Reference | `nodus/grpo-lora` | | Image | `nodus/grpo-lora:2.0.0` | | Task | GRPO | | Modes | Train, Evaluate | | Launcher | Torchrun | | Requires | model yes, data no, environment yes | | Checkpoints | `/nodus/state` (HFTrainer) | ## Parameters [Section titled “Parameters”](#parameters) Set under `spec.parameters`; the server validates them against this schema, and unset parameters take the runtime default. | Parameter | Type | Default | Allowed | Description | | ---------------------------------- | ------- | ---------- | ------------------------------------------------------ | -------------------------------------------------------------------------------------------------------------------------------- | | `attention` | string | none | `eager`, `sdpa`, `flash_attention_2` | Attention kernel | | `batchSize` | integer | `1` | 1 to 1024 | Examples per GPU per step | | `beta` | number | `0` | 0 to 10 | KL penalty toward the base model; 0 disables the reference model | | `distributed.findUnusedParameters` | boolean | none | | DDP: tolerate parameters without gradients | | `distributed.strategy` | string | `Auto` | `Auto`, `DDP`, `FSDP`, `ZeRO1`, `ZeRO2`, `ZeRO3` | Auto and DDP replicate the model; FSDP and ZeRO shard it | | `epochs` | number | `1` | 0.01 to 100 | Passes over the data when `steps` is -1 | | `evalSteps` | integer | none | 1 to 100000 | Steps between validation passes; unset evaluates at the end only | | `evaluationBatchSize` | integer | none | 1 to 1024 | Prompts generated together during Environment baseline and final evaluation | | `generationBatchSize` | integer | none | 2 to 65536 | Total sampled completions per generation batch; must divide evenly across GPUs and numGenerations | | `gradientAccumulation` | integer | `4` | 1 to 1024 | Steps whose gradients are summed before an update | | `gradientCheckpointing` | boolean | `true` | | Recompute activations to save GPU memory | | `learningRate` | number | `0.000001` | 0 to 1 | Peak learning rate | | `loggingSteps` | integer | `10` | 1 to 10000 | Steps between metric reports | | `lora.alpha` | integer | `16` | 1 to 4096 | Adapter scaling numerator | | `lora.dropout` | number | `0.05` | 0 to 0.9 | Adapter dropout | | `lora.r` | integer | `8` | 1 to 1024 | Adapter rank | | `lora.targetModules` | any | none | | Module names, or all-linear | | `lrScheduler` | string | `cosine` | `cosine`, `linear`, `constant`, `constant_with_warmup` | Learning-rate schedule | | `maxCompletionLength` | integer | `64` | 8 to 32768 | Tokens per completion; in a multi-turn episode, tokens per reply | | `maxEpisodeTokens` | integer | none | 8 to 131072 | Tokens of a whole multi-turn episode after the prompt, the model’s and the Environment’s | | `maxGradNorm` | number | none | 0 to 100 | Gradient clipping norm | | `maxLength` | integer | `2048` | 16 to 131072 | Tokens per example after truncation (packed block size for pretraining) | | `maxTurns` | integer | none | 1 to 31 | Model turns per episode of a multi-turn Environment; 1 grades the first reply alone | | `method` | string | `LoRA` | `Full`, `LoRA`, `QLoRA` | Full fine-tuning, a LoRA adapter, or LoRA on a 4-bit quantized base (GPU only) | | `numGenerations` | integer | `4` | 2 to 64 | Completions sampled per prompt (the GRPO group) | | `precision` | string | `Auto` | `Auto`, `bf16`, `fp16`, `fp32` | Numeric precision; Auto is bf16 where the GPU supports it, else fp16 | | `sampling.temperature` | number | `0.7` | 0 to 5 | Sampling temperature | | `sampling.topK` | integer | none | 0 to 1000000 | Keep the highest-probability K tokens; 0 disables top-k filtering | | `sampling.topP` | number | `1` | 0 to 1 | Nucleus sampling mass | | `saveSteps` | integer | `25` | 1 to 100000 | Steps between recovery checkpoints | | `saveTotalLimit` | integer | none | 1 to 10 | Recovery checkpoints kept in the state directory | | `seed` | integer | `42` | 0 to 2147483647 | Seed for data order, splits and initialization | | `steps` | integer | `50` | -1 to 1000000 | Optimizer steps; -1 trains for `epochs` instead | | `taskFilter` | object | none | at most 8 keys | Keep only tasks whose metadata has these values (such as family on reasoning-gym), before trainTasks and heldOutTasks count them | | `warmupRatio` | number | `0.03` | 0 to 0.5 | Fraction of steps spent warming up the learning rate | | `weightDecay` | number | none | 0 to 1 | AdamW weight decay | ## Outputs [Section titled “Outputs”](#outputs) | Output | Path | | --------- | ---------------- | | `outputs` | `/nodus/outputs` | Every run also writes `results.json`, `provenance.json` and a sha256 `manifest.json`. Each file under `/nodus/outputs` is an output of the TrainingJob, named by its path: download one with `nodus cp tj/:outputs/results.json ./results.json`. ## Presets [Section titled “Presets”](#presets) Without `spec.resources`, the TrainingJob takes the resources of the preset matching its model, method and quantization, else the runtime default. | Model | Method | Quantization | GPUs | | --------------- | ------ | ------------ | --------------------------------- | | Qwen/Qwen3-0.6B | LoRA | none | 1 × RTX-4090 or L40S or RTX-A6000 | | Qwen/Qwen3-1.7B | LoRA | none | 1 × RTX-4090 or L40S or RTX-A6000 | | Qwen/Qwen3-4B | QLoRA | nf4 | 1 × RTX-4090 or L40S or RTX-A6000 | | Qwen/Qwen3-4B | LoRA | none | 1 × L40S or RTX-A6000 or A100-80G | | Qwen/Qwen3-8B | QLoRA | nf4 | 1 × L40S or RTX-A6000 or A100-80G | Default resources: 1 × RTX-4090 or L40S or RTX-A6000 or A100-80G. ## Estimates [Section titled “Estimates”](#estimates) The dry-run estimate multiplies these by the run’s steps or tasks; the nightly runs re-measure them. | Accelerator | Seconds | Per | | ----------- | ------- | -------------- | | A100-80G | 0.7 | secondsPerStep | | H100 | 0.45 | secondsPerStep | | L40S | 1 | secondsPerStep | | RTX-4090 | 1.2 | secondsPerStep | | RTX-A6000 | 1.3 | secondsPerStep | ## Example [Section titled “Example”](#example) ```yaml apiVersion: nodus.dev/v1beta1 kind: TrainingJob metadata: {name: my-grpo-lora} spec: runtime: nodus/grpo-lora model: {uri: "hf://Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca"} environment: {name: nodus/graph-coloring@1.0.0, trainTasks: 200, heldOutTasks: 64, seed: 42} maxCostUSD: "5.00" ``` # kto > Kahneman-Tversky optimization on unpaired desirable and undesirable completions Beta TrainingJob and TrainingRuntime are Beta: fields can still change before GA. Kahneman-Tversky optimization on unpaired desirable and undesirable completions. | Field | Value | | ----------- | ----------------------------------- | | Reference | `nodus/kto` | | Image | `nodus/kto:2.0.0` | | Task | KTO | | Modes | Train | | Launcher | Torchrun | | Requires | model yes, data yes, environment no | | Checkpoints | `/nodus/state` (HFTrainer) | ## Parameters [Section titled “Parameters”](#parameters) Set under `spec.parameters`; the server validates them against this schema, and unset parameters take the runtime default. | Parameter | Type | Default | Allowed | Description | | ---------------------------------- | ------- | ---------- | ------------------------------------------------------ | ------------------------------------------------------------------------------ | | `attention` | string | none | `eager`, `sdpa`, `flash_attention_2` | Attention kernel | | `batchSize` | integer | `4` | 1 to 1024 | Examples per GPU per step | | `beta` | number | `0.1` | 0 to 10 | Strength of the KL penalty toward the reference model | | `desirableWeight` | number | `1` | 0 to 100 | Loss weight of desirable examples | | `distributed.findUnusedParameters` | boolean | none | | DDP: tolerate parameters without gradients | | `distributed.strategy` | string | `Auto` | `Auto`, `DDP`, `FSDP`, `ZeRO1`, `ZeRO2`, `ZeRO3` | Auto and DDP replicate the model; FSDP and ZeRO shard it | | `epochs` | number | `1` | 0.01 to 100 | Passes over the data when `steps` is -1 | | `evalSteps` | integer | none | 1 to 100000 | Steps between validation passes; unset evaluates at the end only | | `gradientAccumulation` | integer | `4` | 1 to 1024 | Steps whose gradients are summed before an update | | `gradientCheckpointing` | boolean | `true` | | Recompute activations to save GPU memory | | `learningRate` | number | `0.000005` | 0 to 1 | Peak learning rate | | `loggingSteps` | integer | `10` | 1 to 10000 | Steps between metric reports | | `lora.alpha` | integer | `32` | 1 to 4096 | Adapter scaling numerator | | `lora.dropout` | number | `0.05` | 0 to 0.9 | Adapter dropout | | `lora.r` | integer | `16` | 1 to 1024 | Adapter rank | | `lora.targetModules` | any | none | | Module names, or all-linear | | `lrScheduler` | string | `cosine` | `cosine`, `linear`, `constant`, `constant_with_warmup` | Learning-rate schedule | | `maxGradNorm` | number | none | 0 to 100 | Gradient clipping norm | | `maxLength` | integer | `1024` | 16 to 131072 | Tokens per example after truncation (packed block size for pretraining) | | `method` | string | `LoRA` | `Full`, `LoRA`, `QLoRA` | Full fine-tuning, a LoRA adapter, or LoRA on a 4-bit quantized base (GPU only) | | `precision` | string | `Auto` | `Auto`, `bf16`, `fp16`, `fp32` | Numeric precision; Auto is bf16 where the GPU supports it, else fp16 | | `saveSteps` | integer | `200` | 1 to 100000 | Steps between recovery checkpoints | | `saveTotalLimit` | integer | none | 1 to 10 | Recovery checkpoints kept in the state directory | | `seed` | integer | `42` | 0 to 2147483647 | Seed for data order, splits and initialization | | `steps` | integer | `-1` | -1 to 1000000 | Optimizer steps; -1 trains for `epochs` instead | | `undesirableWeight` | number | `1` | 0 to 100 | Loss weight of undesirable examples | | `warmupRatio` | number | `0.03` | 0 to 0.5 | Fraction of steps spent warming up the learning rate | | `weightDecay` | number | none | 0 to 1 | AdamW weight decay | ## Outputs [Section titled “Outputs”](#outputs) | Output | Path | | --------- | ---------------- | | `outputs` | `/nodus/outputs` | Every run also writes `results.json`, `provenance.json` and a sha256 `manifest.json`. Each file under `/nodus/outputs` is an output of the TrainingJob, named by its path: download one with `nodus cp tj/:outputs/results.json ./results.json`. ## Presets [Section titled “Presets”](#presets) Without `spec.resources`, the TrainingJob takes the resources of the preset matching its model, method and quantization, else the runtime default. | Model | Method | Quantization | GPUs | | --------------- | ------ | ------------ | --------------------------------- | | Qwen/Qwen3-0.6B | LoRA | none | 1 × RTX-4090 or L40S or RTX-A6000 | | Qwen/Qwen3-0.6B | Full | none | 1 × L40S or RTX-A6000 or A100-80G | | Qwen/Qwen3-1.7B | LoRA | none | 1 × RTX-4090 or L40S or RTX-A6000 | | Qwen/Qwen3-1.7B | Full | none | 1 × A100-80G or H100 | | Qwen/Qwen3-4B | QLoRA | nf4 | 1 × RTX-4090 or L40S or RTX-A6000 | | Qwen/Qwen3-4B | LoRA | none | 1 × L40S or RTX-A6000 or A100-80G | | Qwen/Qwen3-8B | QLoRA | nf4 | 1 × L40S or RTX-A6000 or A100-80G | | Qwen/Qwen3-8B | Full | none | 8 × H100 | Default resources: 1 × RTX-4090 or L40S or RTX-A6000 or A100-80G. ## Estimates [Section titled “Estimates”](#estimates) The dry-run estimate multiplies these by the run’s steps or tasks; the nightly runs re-measure them. | Accelerator | Seconds | Per | | ----------- | ------- | -------------- | | A100-80G | 0.7 | secondsPerStep | | H100 | 0.45 | secondsPerStep | | L40S | 1 | secondsPerStep | | RTX-4090 | 1.2 | secondsPerStep | | RTX-A6000 | 1.3 | secondsPerStep | ## Example [Section titled “Example”](#example) ```yaml apiVersion: nodus.dev/v1beta1 kind: TrainingJob metadata: {name: my-kto} spec: runtime: nodus/kto model: {uri: "hf://Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca"} data: {volume: my-dataset, format: JSONL} maxCostUSD: "5.00" ``` # orpo > Odds-ratio preference optimization: SFT and preference alignment in one pass, no reference model Beta TrainingJob and TrainingRuntime are Beta: fields can still change before GA. Odds-ratio preference optimization: SFT and preference alignment in one pass, no reference model. | Field | Value | | ----------- | ----------------------------------- | | Reference | `nodus/orpo` | | Image | `nodus/orpo:2.0.0` | | Task | ORPO | | Modes | Train | | Launcher | Torchrun | | Requires | model yes, data yes, environment no | | Checkpoints | `/nodus/state` (HFTrainer) | ## Parameters [Section titled “Parameters”](#parameters) Set under `spec.parameters`; the server validates them against this schema, and unset parameters take the runtime default. | Parameter | Type | Default | Allowed | Description | | ---------------------------------- | ------- | ---------- | ------------------------------------------------------ | ------------------------------------------------------------------------------ | | `attention` | string | none | `eager`, `sdpa`, `flash_attention_2` | Attention kernel | | `batchSize` | integer | `4` | 1 to 1024 | Examples per GPU per step | | `beta` | number | `0.1` | 0 to 10 | Weight of the odds-ratio term (lambda in the ORPO paper) | | `distributed.findUnusedParameters` | boolean | none | | DDP: tolerate parameters without gradients | | `distributed.strategy` | string | `Auto` | `Auto`, `DDP`, `FSDP`, `ZeRO1`, `ZeRO2`, `ZeRO3` | Auto and DDP replicate the model; FSDP and ZeRO shard it | | `epochs` | number | `1` | 0.01 to 100 | Passes over the data when `steps` is -1 | | `evalSteps` | integer | none | 1 to 100000 | Steps between validation passes; unset evaluates at the end only | | `gradientAccumulation` | integer | `4` | 1 to 1024 | Steps whose gradients are summed before an update | | `gradientCheckpointing` | boolean | `true` | | Recompute activations to save GPU memory | | `learningRate` | number | `0.000008` | 0 to 1 | Peak learning rate | | `loggingSteps` | integer | `10` | 1 to 10000 | Steps between metric reports | | `lora.alpha` | integer | `32` | 1 to 4096 | Adapter scaling numerator | | `lora.dropout` | number | `0.05` | 0 to 0.9 | Adapter dropout | | `lora.r` | integer | `16` | 1 to 1024 | Adapter rank | | `lora.targetModules` | any | none | | Module names, or all-linear | | `lrScheduler` | string | `cosine` | `cosine`, `linear`, `constant`, `constant_with_warmup` | Learning-rate schedule | | `maxGradNorm` | number | none | 0 to 100 | Gradient clipping norm | | `maxLength` | integer | `1024` | 16 to 131072 | Tokens per example after truncation (packed block size for pretraining) | | `method` | string | `LoRA` | `Full`, `LoRA`, `QLoRA` | Full fine-tuning, a LoRA adapter, or LoRA on a 4-bit quantized base (GPU only) | | `precision` | string | `Auto` | `Auto`, `bf16`, `fp16`, `fp32` | Numeric precision; Auto is bf16 where the GPU supports it, else fp16 | | `saveSteps` | integer | `200` | 1 to 100000 | Steps between recovery checkpoints | | `saveTotalLimit` | integer | none | 1 to 10 | Recovery checkpoints kept in the state directory | | `seed` | integer | `42` | 0 to 2147483647 | Seed for data order, splits and initialization | | `steps` | integer | `-1` | -1 to 1000000 | Optimizer steps; -1 trains for `epochs` instead | | `warmupRatio` | number | `0.03` | 0 to 0.5 | Fraction of steps spent warming up the learning rate | | `weightDecay` | number | none | 0 to 1 | AdamW weight decay | ## Outputs [Section titled “Outputs”](#outputs) | Output | Path | | --------- | ---------------- | | `outputs` | `/nodus/outputs` | Every run also writes `results.json`, `provenance.json` and a sha256 `manifest.json`. Each file under `/nodus/outputs` is an output of the TrainingJob, named by its path: download one with `nodus cp tj/:outputs/results.json ./results.json`. ## Presets [Section titled “Presets”](#presets) Without `spec.resources`, the TrainingJob takes the resources of the preset matching its model, method and quantization, else the runtime default. | Model | Method | Quantization | GPUs | | --------------- | ------ | ------------ | --------------------------------- | | Qwen/Qwen3-0.6B | LoRA | none | 1 × RTX-4090 or L40S or RTX-A6000 | | Qwen/Qwen3-0.6B | Full | none | 1 × L40S or RTX-A6000 or A100-80G | | Qwen/Qwen3-1.7B | LoRA | none | 1 × RTX-4090 or L40S or RTX-A6000 | | Qwen/Qwen3-1.7B | Full | none | 1 × A100-80G or H100 | | Qwen/Qwen3-4B | QLoRA | nf4 | 1 × RTX-4090 or L40S or RTX-A6000 | | Qwen/Qwen3-4B | LoRA | none | 1 × L40S or RTX-A6000 or A100-80G | | Qwen/Qwen3-8B | QLoRA | nf4 | 1 × L40S or RTX-A6000 or A100-80G | | Qwen/Qwen3-8B | Full | none | 8 × H100 | Default resources: 1 × RTX-4090 or L40S or RTX-A6000 or A100-80G. ## Estimates [Section titled “Estimates”](#estimates) The dry-run estimate multiplies these by the run’s steps or tasks; the nightly runs re-measure them. | Accelerator | Seconds | Per | | ----------- | ------- | -------------- | | A100-80G | 0.7 | secondsPerStep | | H100 | 0.45 | secondsPerStep | | L40S | 1 | secondsPerStep | | RTX-4090 | 1.2 | secondsPerStep | | RTX-A6000 | 1.3 | secondsPerStep | ## Example [Section titled “Example”](#example) ```yaml apiVersion: nodus.dev/v1beta1 kind: TrainingJob metadata: {name: my-orpo} spec: runtime: nodus/orpo model: {uri: "hf://Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca"} data: {volume: my-dataset, format: JSONL} maxCostUSD: "5.00" ``` # reward-model > Reward-model training: a one-logit head scored on chosen versus rejected Beta TrainingJob and TrainingRuntime are Beta: fields can still change before GA. Reward-model training: a one-logit head scored on chosen versus rejected. | Field | Value | | ----------- | ----------------------------------- | | Reference | `nodus/reward-model` | | Image | `nodus/reward-model:2.0.0` | | Task | Reward | | Modes | Train | | Launcher | Torchrun | | Requires | model yes, data yes, environment no | | Checkpoints | `/nodus/state` (HFTrainer) | ## Parameters [Section titled “Parameters”](#parameters) Set under `spec.parameters`; the server validates them against this schema, and unset parameters take the runtime default. | Parameter | Type | Default | Allowed | Description | | ---------------------------------- | ------- | --------- | ------------------------------------------------------ | ------------------------------------------------------------------------------ | | `attention` | string | none | `eager`, `sdpa`, `flash_attention_2` | Attention kernel | | `batchSize` | integer | `4` | 1 to 1024 | Examples per GPU per step | | `distributed.findUnusedParameters` | boolean | none | | DDP: tolerate parameters without gradients | | `distributed.strategy` | string | `Auto` | `Auto`, `DDP`, `FSDP`, `ZeRO1`, `ZeRO2`, `ZeRO3` | Auto and DDP replicate the model; FSDP and ZeRO shard it | | `epochs` | number | `1` | 0.01 to 100 | Passes over the data when `steps` is -1 | | `evalSteps` | integer | none | 1 to 100000 | Steps between validation passes; unset evaluates at the end only | | `gradientAccumulation` | integer | `4` | 1 to 1024 | Steps whose gradients are summed before an update | | `gradientCheckpointing` | boolean | `true` | | Recompute activations to save GPU memory | | `learningRate` | number | `0.00001` | 0 to 1 | Peak learning rate | | `loggingSteps` | integer | `10` | 1 to 10000 | Steps between metric reports | | `lora.alpha` | integer | `32` | 1 to 4096 | Adapter scaling numerator | | `lora.dropout` | number | `0.05` | 0 to 0.9 | Adapter dropout | | `lora.r` | integer | `16` | 1 to 1024 | Adapter rank | | `lora.targetModules` | any | none | | Module names, or all-linear | | `lrScheduler` | string | `cosine` | `cosine`, `linear`, `constant`, `constant_with_warmup` | Learning-rate schedule | | `maxGradNorm` | number | none | 0 to 100 | Gradient clipping norm | | `maxLength` | integer | `1024` | 16 to 131072 | Tokens per example after truncation (packed block size for pretraining) | | `method` | string | `LoRA` | `Full`, `LoRA`, `QLoRA` | Full fine-tuning, a LoRA adapter, or LoRA on a 4-bit quantized base (GPU only) | | `precision` | string | `Auto` | `Auto`, `bf16`, `fp16`, `fp32` | Numeric precision; Auto is bf16 where the GPU supports it, else fp16 | | `saveSteps` | integer | `200` | 1 to 100000 | Steps between recovery checkpoints | | `saveTotalLimit` | integer | none | 1 to 10 | Recovery checkpoints kept in the state directory | | `seed` | integer | `42` | 0 to 2147483647 | Seed for data order, splits and initialization | | `steps` | integer | `-1` | -1 to 1000000 | Optimizer steps; -1 trains for `epochs` instead | | `warmupRatio` | number | `0.03` | 0 to 0.5 | Fraction of steps spent warming up the learning rate | | `weightDecay` | number | none | 0 to 1 | AdamW weight decay | ## Outputs [Section titled “Outputs”](#outputs) | Output | Path | | --------- | ---------------- | | `outputs` | `/nodus/outputs` | Every run also writes `results.json`, `provenance.json` and a sha256 `manifest.json`. Each file under `/nodus/outputs` is an output of the TrainingJob, named by its path: download one with `nodus cp tj/:outputs/results.json ./results.json`. ## Presets [Section titled “Presets”](#presets) Without `spec.resources`, the TrainingJob takes the resources of the preset matching its model, method and quantization, else the runtime default. | Model | Method | Quantization | GPUs | | --------------- | ------ | ------------ | --------------------------------- | | Qwen/Qwen3-0.6B | LoRA | none | 1 × RTX-4090 or L40S or RTX-A6000 | | Qwen/Qwen3-0.6B | Full | none | 1 × L40S or RTX-A6000 or A100-80G | | Qwen/Qwen3-1.7B | LoRA | none | 1 × RTX-4090 or L40S or RTX-A6000 | | Qwen/Qwen3-1.7B | Full | none | 1 × A100-80G or H100 | | Qwen/Qwen3-4B | QLoRA | nf4 | 1 × RTX-4090 or L40S or RTX-A6000 | | Qwen/Qwen3-4B | LoRA | none | 1 × L40S or RTX-A6000 or A100-80G | | Qwen/Qwen3-8B | QLoRA | nf4 | 1 × L40S or RTX-A6000 or A100-80G | | Qwen/Qwen3-8B | Full | none | 8 × H100 | Default resources: 1 × RTX-4090 or L40S or RTX-A6000 or A100-80G. ## Estimates [Section titled “Estimates”](#estimates) The dry-run estimate multiplies these by the run’s steps or tasks; the nightly runs re-measure them. | Accelerator | Seconds | Per | | ----------- | ------- | -------------- | | A100-80G | 0.7 | secondsPerStep | | H100 | 0.45 | secondsPerStep | | L40S | 1 | secondsPerStep | | RTX-4090 | 1.2 | secondsPerStep | | RTX-A6000 | 1.3 | secondsPerStep | ## Example [Section titled “Example”](#example) ```yaml apiVersion: nodus.dev/v1beta1 kind: TrainingJob metadata: {name: my-reward-model} spec: runtime: nodus/reward-model model: {uri: "hf://Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca"} data: {volume: my-dataset, format: JSONL} maxCostUSD: "5.00" ``` # sft > Supervised fine-tuning (full, LoRA, QLoRA) and pretraining with task: Pretrain Beta TrainingJob and TrainingRuntime are Beta: fields can still change before GA. Supervised fine-tuning (full, LoRA, QLoRA) and pretraining with task: Pretrain. | Field | Value | | ----------- | ----------------------------------- | | Reference | `nodus/sft` | | Image | `nodus/sft:2.0.0` | | Task | SFT | | Modes | Train | | Launcher | Torchrun | | Requires | model yes, data yes, environment no | | Checkpoints | `/nodus/state` (HFTrainer) | ## Parameters [Section titled “Parameters”](#parameters) Set under `spec.parameters`; the server validates them against this schema, and unset parameters take the runtime default. | Parameter | Type | Default | Allowed | Description | | ---------------------------------- | ------- | ----------- | ------------------------------------------------------ | ----------------------------------------------------------------------------------- | | `architecture` | object | none | | Scratch only: config overrides such as num_hidden_layers or hidden_size | | `assistantOnlyLoss` | boolean | none | | Conversational data: train on assistant turns only | | `attention` | string | none | `eager`, `sdpa`, `flash_attention_2` | Attention kernel | | `batchSize` | integer | `4` | 1 to 1024 | Examples per GPU per step | | `completionOnlyLoss` | boolean | none | | Prompt/completion data: train on the completion tokens only | | `dataProcesses` | integer | none | 1 to 64 | Processes for tokenizing the data | | `distributed.findUnusedParameters` | boolean | none | | DDP: tolerate parameters without gradients | | `distributed.strategy` | string | `Auto` | `Auto`, `DDP`, `FSDP`, `ZeRO1`, `ZeRO2`, `ZeRO3` | Auto and DDP replicate the model; FSDP and ZeRO shard it | | `epochs` | number | `1` | 0.01 to 100 | Passes over the data when `steps` is -1 | | `evalSteps` | integer | none | 1 to 100000 | Steps between validation passes; unset evaluates at the end only | | `gradientAccumulation` | integer | `4` | 1 to 1024 | Steps whose gradients are summed before an update | | `gradientCheckpointing` | boolean | `true` | | Recompute activations to save GPU memory | | `initialization` | string | `Continued` | `Continued`, `Scratch` | Pretraining: continue from the model’s weights or start from a fresh initialization | | `learningRate` | number | `0.0002` | 0 to 1 | Peak learning rate | | `loggingSteps` | integer | `10` | 1 to 10000 | Steps between metric reports | | `lora.alpha` | integer | `32` | 1 to 4096 | Adapter scaling numerator | | `lora.dropout` | number | `0.05` | 0 to 0.9 | Adapter dropout | | `lora.r` | integer | `16` | 1 to 1024 | Adapter rank | | `lora.targetModules` | any | none | | Module names, or all-linear | | `lrScheduler` | string | `cosine` | `cosine`, `linear`, `constant`, `constant_with_warmup` | Learning-rate schedule | | `maxGradNorm` | number | none | 0 to 100 | Gradient clipping norm | | `maxLength` | integer | `2048` | 16 to 131072 | Tokens per example after truncation (packed block size for pretraining) | | `method` | string | `LoRA` | `Full`, `LoRA`, `QLoRA` | Full fine-tuning, a LoRA adapter, or LoRA on a 4-bit quantized base (GPU only) | | `packing` | boolean | none | | Pack short examples into full-length blocks (default on for pretraining) | | `precision` | string | `Auto` | `Auto`, `bf16`, `fp16`, `fp32` | Numeric precision; Auto is bf16 where the GPU supports it, else fp16 | | `saveSteps` | integer | `200` | 1 to 100000 | Steps between recovery checkpoints | | `saveTotalLimit` | integer | none | 1 to 10 | Recovery checkpoints kept in the state directory | | `seed` | integer | `42` | 0 to 2147483647 | Seed for data order, splits and initialization | | `steps` | integer | `-1` | -1 to 1000000 | Optimizer steps; -1 trains for `epochs` instead | | `warmupRatio` | number | `0.03` | 0 to 0.5 | Fraction of steps spent warming up the learning rate | | `weightDecay` | number | none | 0 to 1 | AdamW weight decay | ## Outputs [Section titled “Outputs”](#outputs) | Output | Path | | --------- | ---------------- | | `outputs` | `/nodus/outputs` | Every run also writes `results.json`, `provenance.json` and a sha256 `manifest.json`. Each file under `/nodus/outputs` is an output of the TrainingJob, named by its path: download one with `nodus cp tj/:outputs/results.json ./results.json`. ## Presets [Section titled “Presets”](#presets) Without `spec.resources`, the TrainingJob takes the resources of the preset matching its model, method and quantization, else the runtime default. | Model | Method | Quantization | GPUs | | --------------- | ------ | ------------ | --------------------------------- | | Qwen/Qwen3-0.6B | LoRA | none | 1 × RTX-4090 or L40S or RTX-A6000 | | Qwen/Qwen3-0.6B | Full | none | 1 × L40S or RTX-A6000 or A100-80G | | Qwen/Qwen3-1.7B | LoRA | none | 1 × RTX-4090 or L40S or RTX-A6000 | | Qwen/Qwen3-1.7B | Full | none | 1 × A100-80G or H100 | | Qwen/Qwen3-4B | QLoRA | nf4 | 1 × RTX-4090 or L40S or RTX-A6000 | | Qwen/Qwen3-4B | LoRA | none | 1 × L40S or RTX-A6000 or A100-80G | | Qwen/Qwen3-8B | QLoRA | nf4 | 1 × L40S or RTX-A6000 or A100-80G | | Qwen/Qwen3-8B | Full | none | 8 × H100 | Default resources: 1 × RTX-4090 or L40S or RTX-A6000 or A100-80G. ## Estimates [Section titled “Estimates”](#estimates) The dry-run estimate multiplies these by the run’s steps or tasks; the nightly runs re-measure them. | Accelerator | Seconds | Per | | ----------- | ------- | -------------- | | A100-80G | 0.7 | secondsPerStep | | H100 | 0.45 | secondsPerStep | | L40S | 1 | secondsPerStep | | RTX-4090 | 1.2 | secondsPerStep | | RTX-A6000 | 1.3 | secondsPerStep | ## Example [Section titled “Example”](#example) ```yaml apiVersion: nodus.dev/v1beta1 kind: TrainingJob metadata: {name: my-sft} spec: runtime: nodus/sft model: {uri: "hf://Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca"} data: {volume: my-dataset, format: JSONL} maxCostUSD: "5.00" ``` # Server configuration > Every environment variable nodus-server, the node agent and the CLI read. Nodus processes read `NODUS_`-prefixed environment variables, validated at startup: a missing or invalid value stops the process with an error that names the variable. Secrets come from the environment only. ## nodus-server (roles api, controller, gateway, inference) [Section titled “nodus-server (roles api, controller, gateway, inference)”](#nodus-server-roles-api-controller-gateway-inference) | Variable | Default | Required | Description | | ------------------------------------ | ------------------------------------------ | -------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `NODUS_ASSISTANT_DOCS_INDEX_URL` | `https://nodus-compute.ai/docs/index.json` | no | The site’s runtime docs index, which the assistant searches and cites; the api role caches it for 15 minutes. Roles: api. | | `NODUS_AUTH_ISSUER` | — | no | Expected iss claim of user access tokens (https\://.supabase.co/auth/v1); required outside NODUS_ENV=local, where empty skips the issuer check. Roles: api, gateway. | | `NODUS_AUTH_JWKS_URL` | — | no | JWKS URL of the Auth signing keys; empty derives \/.well-known/jwks.json. Roles: api, gateway. | | `NODUS_AUTH_URL` | — | no | Supabase Auth base URL (https\://.supabase.co/auth/v1; on compose); empty, which only NODUS_ENV=local allows, accepts only API keys and ServiceAccount tokens. Roles: api, gateway. | | `NODUS_CORS_ORIGINS` | — | no | Comma-separated browser origins allowed to call the API: the console and admin apps (ADR-107). Roles: api, gateway. | | `NODUS_DATABASE_MAX_CONNS` | `0` | no | Pool size override; 0 uses the role budget of ADR-070 (api 8, controller 12, gateway 6, inference 4). | | `NODUS_DATABASE_URL` | — | yes | Postgres URL: Supavisor session mode (port 5432) in the cloud, the compose postgres locally. | | `NODUS_ENV` | `local` | no | Deployment: local (compose, CI, tests), staging or prod. | | `NODUS_GITHUB_APP` | — | no | GitHub App as JSON: app_id, slug, private_key (PEM), webhook_secret, client_id, client_secret; empty disables GitHub Connections. Roles: api, controller, gateway. | | `NODUS_HTTP_ADDR` | `:8080` | no | Public HTTP listener: api, gateway client streams, inference and health. | | `NODUS_INFERENCE_CONSOLE_ORIGINS` | — | no | Comma-separated browser origins allowed to call the data plane with a user token (the console playground). Roles: inference. | | `NODUS_INFERENCE_REPLICAS` | `1` | no | Fleet size of the inference role; the per-replica rate limits divide by it. Roles: inference. | | `NODUS_INFERENCE_UPSTREAMS` | — | no | Inference upstreams as JSON: groq_keys (list, at most 8), groq_base_url, wafer_key, wafer_base_url, openrelay_key, openrelay_base_url, openrelay_catalog_url, anthropic_key, anthropic_base_url, jev_key, jev_endpoint, jev_model; every field optional, URLs https; empty serves no upstream unless its dedicated key is configured. Roles: controller, inference. | | `NODUS_INTERNAL_ADDR` | `:9090` | no | Internal stream-bridge listener. Roles: gateway. | | `NODUS_KMS_KEY_URI` | — | no | Key-encryption key: aws-kms\://arn:aws:kms:… in the cloud, local://\ locally; empty uses a fixed development key when NODUS_ENV=local. | | `NODUS_LOCAL_DOCKER_HOST` | — | no | Docker endpoint the local provider creates node containers on, such as unix:///Users/me/.colima/default/docker.sock; empty keeps providers.yaml’s or unix:///var/run/docker.sock. Roles: controller. | | `NODUS_LOCAL_NODE_BINDS` | — | no | Comma-separated Docker binds of every local node container, such as /path/to/worktree:/src. Roles: controller. | | `NODUS_LOCAL_NODE_COMMAND` | — | no | Shell command a local node container runs (sh -c); empty keeps providers.yaml’s or the image’s default. Roles: controller. | | `NODUS_LOCAL_NODE_ENV` | — | no | Comma-separated KEY=value pairs added to every local node container, such as NODUS_GATEWAY_URL=. Roles: controller. | | `NODUS_LOCAL_NODE_IMAGE` | — | no | Image of the local provider’s node containers; empty keeps providers.yaml’s or the compose-built nodus-platform-nodusd. Roles: controller. | | `NODUS_LOCAL_NODE_NETWORK` | — | no | Docker network the local provider’s node containers join; empty keeps providers.yaml’s or nodus-platform_default. Roles: controller. | | `NODUS_LOCAL_NODE_OBJSTORE_ENDPOINT` | — | no | Object store endpoint written into local nodes’ storage grants when they reach the store at another address than NODUS_OBJSTORE_ENDPOINT, such as for a host-run server; empty keeps NODUS_OBJSTORE_ENDPOINT. Roles: controller. | | `NODUS_LOG_LEVEL` | `info` | no | Minimum log level: debug, info, warn or error. | | `NODUS_METRICS_ADDR` | `:9100` | no | Prometheus metrics listener. | | `NODUS_NODE_ADDR` | `:8443` | no | Node protocol listener. Roles: gateway. | | `NODUS_NODUSD_VERSION` | — | no | nodusd release new nodes fetch, such as git-0123456789ab: the server reads nodusd///nodusd.sha256 from the releases bucket and presigns the binary per create. Required outside local. Roles: controller. | | `NODUS_NOTIFY_CONSOLE_URL` | `http://localhost:5173` | no | Console base URL that links in emails point at. Roles: api, controller. | | `NODUS_NOTIFY_FROM` | — | no | Sender of every notice as an RFC 5322 address, such as Nodus \; required with NODUS_RESEND_API_KEY. Roles: api, controller. | | `NODUS_OBJSTORE_ACCESS_KEY_ID` | — | no | Access key id; empty uses the AWS default credential chain. | | `NODUS_OBJSTORE_BUCKET_PREFIX` | — | yes | Bucket name prefix; buckets are -data, -logs, -ephemeral, -registry and -releases. | | `NODUS_OBJSTORE_ENDPOINT` | — | no | S3 API endpoint; empty uses AWS S3. R2: https\://.r2.cloudflarestorage.com. | | `NODUS_OBJSTORE_R2_ACCOUNT_ID` | — | no | Cloudflare account id; set to issue R2 temporary credentials instead of STS. | | `NODUS_OBJSTORE_R2_API_TOKEN` | — | no | Cloudflare API token allowed to create R2 temporary credentials. | | `NODUS_OBJSTORE_REGION` | `auto` | no | S3 signing region; R2 uses auto. | | `NODUS_OBJSTORE_SECRET_ACCESS_KEY` | — | no | Secret access key. | | `NODUS_OPENRELAY_API_KEY` | — | no | Dedicated OpenRelay inference key. When nonempty, overrides only openrelay_key in NODUS_INFERENCE_UPSTREAMS; empty preserves the composite key and all other upstream settings. Roles: controller, inference. | | `NODUS_OPS_ACCESS_AUDIENCE` | — | no | Deprecated; accepted for rollout compatibility and unused by admin authentication. Roles: api. | | `NODUS_OPS_ACCESS_TEAM_DOMAIN` | — | no | Deprecated; accepted for rollout compatibility and unused by admin authentication. Roles: api. | | `NODUS_OPS_ADDR` | `:8081` | no | Ops API listener. Roles: api. | | `NODUS_OPS_CONSOLE_ORIGIN` | — | no | Admin console origin (.) allowed to call the ops API from the browser; empty allows no cross-origin caller. Roles: api. | | `NODUS_OPS_SUPABASE_ISSUER` | — | no | Supabase Auth issuer (https\://.supabase.co/auth/v1) whose verified sessions are checked against the fixed admin email allowlist; empty refuses every ops request. Roles: api. | | `NODUS_POOLS` | — | no | BYOC pools and cloud accounts as JSON: install_script_url (./install/nodusd.sh), node_gateway_url (wss\://nodes., the ADR-123 carrier), api_url and inference_url (workload proxy origins; default to deployed release origins), agent_url (ARCH becomes amd64 or arm64), agent_sha256 ({arch: sha256}), agent_signer_identity (cosign certificate identity regexp), aws_observer_role_arn, aws_template_url, google_client_id, google_client_secret, google_redirect_url, consent_key (32+ characters), console_url. The installer keys replace the one the API otherwise serves from the deployed nodusd release (NODUS_NODUSD_VERSION); without either, EnrollmentTokens are refused, and without a cloud’s keys its onboarding is off. Roles: controller. | | `NODUS_POSTHOG_READ_API_KEY` | — | no | Project-scoped PostHog query:read credential for internal aggregate reports; empty leaves analytics unavailable. Roles: api. | | `NODUS_PROVIDER_SECRETS` | — | no | Where provider account secrets live: ssm:/nodus-platform//app/providers/ (SSM SecureString parameters, ADR-122) or secretsmanager: (a bare prefix also means Secrets Manager); each account’s secretRef is appended. Required outside local. Roles: controller. | | `NODUS_PUBLIC_API_URL` | — | no | Public origin of the API (.), which hosted MCP’s protected-resource metadata and 401 challenge name (RFC 9728); empty uses : under NODUS_ENV=local and leaves /mcp unmounted elsewhere. Roles: api, gateway. | | `NODUS_REGISTRY_HOST` | — | no | Host of the Nodus registry, registry. (localhost:5000 in compose); empty disables image builds, mirrors and registry tokens. Roles: api, controller. | | `NODUS_REGISTRY_SIGNING_KEY` | — | no | P-256 private key, PEM (SEC 1 or PKCS #8), that signs registry bearer tokens; the registry trusts its public key. Roles: api, controller. | | `NODUS_RESEND_API_KEY` | — | no | Resend API key; empty logs each email’s template and recipient count instead of sending, which only NODUS_ENV=local allows. Roles: api, controller. | | `NODUS_SENTRY_DSN` | — | no | Sentry DSN; empty disables error reporting. | | `NODUS_STRIPE_API_BASE` | — | no | Stripe API base URL override; the compose stack points it at stripe-mock. Empty uses api.stripe.com. | | `NODUS_STRIPE_LIVE` | `false` | no | Whether this deployment takes live-mode payments; events whose livemode differs are rejected. | | `NODUS_STRIPE_SECRET_KEY` | — | no | Stripe secret or restricted key (sk_test\_/rk_test\_ outside prod), agreeing with NODUS_STRIPE_LIVE; empty turns payments off. | | `NODUS_STRIPE_WEBHOOK_SECRET` | — | no | Signing secret (whsec\_…) of the POST /webhooks/stripe endpoint; empty refuses every delivery. | | `NODUS_STRIPE_WEBHOOK_URL` | — | no | Public URL of POST /webhooks/stripe that nodus-server seed registers as the Stripe webhook endpoint; empty leaves the endpoint alone. | | `NODUS_TOKEN_PEPPERS` | — | no | Token HMAC peppers as :, comma-separated, KMS-encrypted in the cloud; the highest version signs new tokens. Empty uses a fixed development pepper when NODUS_ENV=local. | | `NODUS_USERCONTENT_DOMAIN` | — | no | Registrable domain of preview URLs and Workspace browser tools, https\://-. (ADR-020); never the API’s domain, and one Nodus owns, since it receives each exchange token. Unset outside NODUS_ENV=local turns the browser tools and previews off; NODUS_ENV=local defaults it to nodus-usercontent.net. Roles: controller, gateway. | | `NODUS_WATCH_DATABASE_URL` | — | no | Postgres URL of the watch tailer’s one connection, logged in as nodus_watch (data-model §12); empty uses NODUS_DATABASE_URL, whose login must then bypass row-level security (compose). Roles: api, gateway. | | `NODUS_WEBHOOKS_ALLOW_PREFIXES` | — | no | Address prefixes the SSRF dialer would refuse but deliveries may reach, comma-separated; only NODUS_ENV=local allows any (compose receivers). Roles: controller. | | `NODUS_WORKSPACE_EDGE` | — | no | Optional JSON browser ingress: domain (account.workers.dev), backend_origin (dedicated HTTPS gateway origin), key (32 random bytes, standard base64). Shared only by controller and gateway. Roles: controller, gateway. | | `NODUS_WORKSPACE_EDGE_ACCOUNT` | — | no | Optional JSON browser endpoint provisioning: account_id and api_token with Workers Scripts edit. Controller only; required with WORKSPACE_EDGE. Roles: controller. | ## nodusd (node agent) [Section titled “nodusd (node agent)”](#nodusd-node-agent) | Variable | Default | Required | Description | | -------------------------- | --------------------------------- | -------- | ---------------------------------------------------------------------- | | `NODUS_GATEWAY_URL` | — | yes | Node protocol endpoint the agent dials, such as .. | | `NODUS_STATE_DIR` | `/var/lib/nodus` | no | Agent state directory: credentials, ring buffers, re-adoption records. | | `NODUS_CONTAINERD_ADDRESS` | `/run/containerd/containerd.sock` | no | containerd socket. | | `NODUS_LOG_LEVEL` | `info` | no | Minimum log level: debug, info, warn or error. | | `NODUS_SENTRY_DSN` | — | no | Sentry DSN; empty disables error reporting. | ## nodus CLI and SDK (client environment, ADR-069) [Section titled “nodus CLI and SDK (client environment, ADR-069)”](#nodus-cli-and-sdk-client-environment-adr-069) | Variable | Default | Required | Description | | ---------------- | ------- | -------- | --------------------------------------------------------------------------------- | | `NODUS_API_KEY` | — | no | API key (nodus_sk\_…); overrides the stored login. | | `NODUS_API_URL` | — | no | API base URL, such as .. | | `NODUS_BASE_URL` | — | no | Deprecated alias of NODUS_API_URL, accepted with a warning for one major version. | | `NODUS_ORG` | — | no | Organization name or id. | | `NODUS_PROJECT` | — | no | Project name. | | `NODUS_CONTEXT` | — | no | Named context from the config file. | | `NODUS_CONFIG` | — | no | Config file path. | # Troubleshooting > What to check when sign-in, a launch, a running job, outputs or billing do not behave as you expect. Start with the object itself. `nodus describe /` shows its phase, conditions, recent events, attempts and cost so far, and every error names a code with a fix and a docs link. When you contact support, include the `requestId` from the error. ## Signing in [Section titled “Signing in”](#signing-in) * **The browser never opens.** Run `nodus login --device` and approve the code from any device. * **A CI machine needs a key.** Create an API key in the console (or a ServiceAccount key) and pipe it in: `echo "$NODUS_API_KEY" | nodus login --with-token`. `NODUS_API_KEY` alone also works. * **The wrong org or project.** `nodus whoami` prints the principal, org, project and scopes in use. `nodus config get-contexts` lists every org you signed in to and `nodus config use-context` switches between them. ## A launch is refused [Section titled “A launch is refused”](#a-launch-is-refused) | Code | Meaning | What to do | | --------------------------- | ------------------------------------------------------------- | ---------------------------------------------------------------------- | | `InsufficientCredits` (402) | The hold for this launch is larger than your available credit | Top up with `nodus billing top-up `, or lower `--max-cost` | | `BudgetExceeded` (402) | A budget that covers this object would be exceeded | Raise the budget or launch in a project it does not cover | | `QuotaExceeded` (429) | An org or project quota is reached | `nodus get quotas`; starter limits lift at your first purchase | | `CapacityUnavailable` (503) | No offering matches the requirements right now | Allow more accelerators or regions, or wait for the ETA in the message | | `Invalid` (422) | The spec fails validation | The message names each field; `nodus explain job.spec` documents them | Every code has its own page under [Error codes](/docs/reference/errors/). ## A job does not start [Section titled “A job does not start”](#a-job-does-not-start) * **`Pending` for a long time.** `nodus describe job/` shows why under Conditions and Events, including the offerings considered and why others were rejected. * **The image fails to pull.** Check the image reference and, for a private registry, the pull Secret it names. ## A job fails or stops [Section titled “A job fails or stops”](#a-job-fails-or-stops) * **The command exits non-zero.** `nodus logs job/` shows your program’s output; the phase is `Failed` with the exit code. `nodus run` exits with the same code. * **It stopped when credit ran out.** A running job stops gracefully inside its reserved amount and reports it in its conditions. Add credit and it resumes from its last checkpoint. * **The CLI itself failed.** Exit code `125` is a Nodus or API error and `124` a timeout, never your program’s. ## Outputs and checkpoints [Section titled “Outputs and checkpoints”](#outputs-and-checkpoints) * **An output is missing.** Only paths declared with `--output NAME=/path` (or `spec.outputs`) are collected. List them with `nodus get job/ -o jsonpath='{.status.outputs}'`. * **A resumed job started from scratch.** Nodus restores the files your program saved in its checkpoint directory (`NODUS_CHECKPOINT_DIR`); your program must load them on start. Restoring files does not restore process memory. ## Billing [Section titled “Billing”](#billing) * **A charge looks higher than the run time.** Rented machines bill from creation to confirmed deletion, including start-up and shutdown. `nodus get usage --field-selector object.name= --group-by segment` itemizes it. * **Where did my credit go?** `nodus billing` shows available credit, open holds and budgets; `nodus get usage --group-by project` breaks spending down.