# Run code

---

Jobs, interactive development, sandboxes and remote Python.

---

# Which Nodus compute feature should I use?

> Compare Jobs, Workspaces, Sandboxes and Functions. Choose the right way to run code, call models, run agents or train a model on Nodus.

Source: https://nodus-platform-site.pages.dev/docs/guides/choose-compute/
Build revision: 211ad9f836655b1c3a2668c4693e442471f28614

Use a **Job** when you have a command that should run to completion. Use a **Workspace** for interactive GPU or CPU development, a **Sandbox** for isolated CPU code execution, and a **Function** to call Python remotely or in parallel. Nodus also provides hosted inference, managed agents and training workflows.

## Choose by the work you need to do

|I need to…|Start with|Why it fits|
|-|-|-|
|Run a script, batch task or existing training command|[Jobs](https://nodus-platform-site.pages.dev/docs/guides/jobs/)|Runs a command to completion with logs, outputs and resource and cost limits|
|Develop with SSH, VS Code or JupyterLab|[Workspaces](https://nodus-platform-site.pages.dev/docs/guides/workspaces/)|An interactive GPU or CPU machine with a saved home directory|
|Execute agent-generated or untrusted code|[Sandboxes](https://nodus-platform-site.pages.dev/docs/guides/sandboxes/)|An isolated CPU container driven by command and file requests, with idle stopping|
|Call Python remotely or map it over many inputs|[Functions](https://nodus-platform-site.pages.dev/docs/guides/functions/)|Remote calls with workers that scale with demand|
|Call a hosted language model|[Inference (Beta)](https://nodus-platform-site.pages.dev/docs/guides/inference/)|An OpenAI-compatible model API billed per token|
|Run an agent conversation with shell tools|[Agents (Beta)](https://nodus-platform-site.pages.dev/docs/guides/agents/)|An AgentRun with its own sandbox and recorded execution history|
|Fine-tune or train through managed runtimes|[Training (Beta)](https://nodus-platform-site.pages.dev/docs/guides/training/)|A TrainingJob defines the model, data, runtime and training parameters|

For a first run, follow [your first job](https://nodus-platform-site.pages.dev/docs/getting-started/). It walks through signing in, running a command, following logs and downloading results.

## What is the difference between a Job and a Function?

A Job runs an executable command and finishes when that command exits. A Function is a Python callable deployed in an App; your code invokes it with `.remote()`, `.map()` or `.spawn()`. Choose Jobs for an existing script or batch process. Choose Functions when remote calls should be part of your Python application.

Both can use GPU or CPU workers. Function workers can stay warm between calls; read [Function billing](https://nodus-platform-site.pages.dev/docs/guides/functions/billing/) before choosing idle and scaling settings.

## What is the difference between a Workspace and a Sandbox?

A Workspace is for interactive development with SSH, VS Code or JupyterLab on GPU or CPU compute. Its home directory is backed by a Volume. A Sandbox is an isolated CPU container for programmatic commands and file operations, including code produced by an agent. Its network is closed unless you open it, and it can stop after an idle period.

Read [Workspace storage](https://nodus-platform-site.pages.dev/docs/guides/workspaces/) and [Sandbox isolation](https://nodus-platform-site.pages.dev/docs/concepts/sandboxes-isolation/) before deciding what state and access your task needs.

## Do I need Agents to use Claude Code, Codex or Cursor?

No. Connect your existing coding agent to Nodus through [MCP](https://nodus-platform-site.pages.dev/docs/guides/mcp/) using the [client setup guide](https://nodus-platform-site.pages.dev/connect/). It can work with Nodus resources through that connection. The Agents feature is for running an agent conversation inside Nodus itself.

## What survives an interruption?

Recovery depends on the feature and its configuration. For checkpointed Jobs, your program writes and loads its own files in `NODUS_CHECKPOINT_DIR`; Nodus saves and restores that directory. This does not restore arbitrary process or GPU memory. Restartable Jobs start over, while ephemeral Jobs keep no state.

Read [Checkpoints](https://nodus-platform-site.pages.dev/docs/guides/checkpoints/) for application recovery, [Volumes](https://nodus-platform-site.pages.dev/docs/guides/volumes/) for persistent files and [Outputs](https://nodus-platform-site.pages.dev/docs/guides/outputs/) for results you need to download.

## How do I control spending?

Check the estimate before starting work, set the resource’s supported cost and time limits, and configure [project budgets](https://nodus-platform-site.pages.dev/docs/guides/billing/budgets/). Compute usage, saved storage and model tokens have different billing rules; the [billing guide](https://nodus-platform-site.pages.dev/docs/guides/billing/) explains them. Use the [pricing reference](https://nodus-platform-site.pages.dev/docs/reference/pricing/) for published amounts instead of copying prices from an old example.

## Where should an agent look up exact syntax?

Use the [CLI reference](https://nodus-platform-site.pages.dev/docs/reference/cli/) for commands, [Python SDK reference](https://nodus-platform-site.pages.dev/docs/reference/python/) for signatures and [OpenAPI](https://nodus-platform-site.pages.dev/docs/openapi.json) for v1 request and response fields. Beta resources use the [v1beta1 contract](https://nodus-platform-site.pages.dev/docs/openapi-v1beta1.json).

Every authored guide has a Markdown version with its examples. The [agent guide](https://nodus-platform-site.pages.dev/docs/for-agents/) links a lightweight page manifest and focused topic bundles for retrieval.


---

# Functions

> Run Python functions on Nodus workers, deploy them as an App that stays up, and look them up from anywhere.

Source: https://nodus-platform-site.pages.dev/docs/guides/functions/
Build revision: 211ad9f836655b1c3a2668c4693e442471f28614

A Function is a Python function that runs on Nodus workers. You decorate it, call it from your own code with `.remote()`, `.map()` or `.spawn()`, and Nodus starts workers when calls arrive, keeps them warm for a while and scales them back down. The pages of this guide cover [calling Functions](https://nodus-platform-site.pages.dev/docs/guides/functions/calls), [scaling them](https://nodus-platform-site.pages.dev/docs/guides/functions/scaling) and [what is billed while a worker waits](https://nodus-platform-site.pages.dev/docs/guides/functions/billing).

## Run an App

An App is the group of Functions in one file. `nodus run` creates an ephemeral App, runs your `main` on your machine and deletes the App when `main` returns. Every `.remote()` call runs on a Nodus worker.

examples/functions/map/app.py

```python
"""One Function called three ways, then mapped over a thousand inputs.

Run it with `nodus run examples/functions/map/app.py`.
"""

import nodus

app = nodus.App("fn-map")


@app.function(cpu=1, memory="1Gi", max_workers=4, target_concurrency=8, max_cost=1)
def square(x: int) -> int:
    return x * x


@app.function(cpu=1, memory="1Gi", max_cost=1)
def divide(a: int, b: int) -> float:
    return a / b


@app.local_entrypoint()
def main(n: int = 1000) -> None:
    print("remote:", square.remote(7))  # one call; blocks for the result
    call = square.spawn(8)  # starts a call and returns a handle
    print("spawned:", call.get(timeout=600))
    results = list(square.map(range(n)))  # one call per input, results in input order
    print("map:", len(results), "ordered:", results == [i * i for i in range(n)])
    try:
        divide.remote(1, 0)
    except ZeroDivisionError as exc:  # the remote exception comes back as its own type
        print("raised:", type(exc).__name__)
```

Terminal window

```console
$ nodus run examples/functions/map/app.py
remote: 49
spawned: 64
map: 1000 ordered: True
raised: ZeroDivisionError
```

The directory of the file (minus what `.gitignore` and `.nodusignore` exclude) is uploaded once per content hash, so the workers import the same code you ran. The App renews itself while `main` runs and is deleted shortly after `main` stops, even when your machine loses its connection.

## Deploy an App

`nodus deploy` keeps the App. Its Functions stay available after your program exits, and any program can call them.

examples/functions/warm-pool/app.py

```python
"""A Function that keeps one worker warm, so a call never waits for a start.

Deploy it with `nodus deploy examples/functions/warm-pool/app.py`. The idle worker is billed at the Function's
worker rate for as long as it stays warm; set `min_workers=0` to pay only while calls run.
"""

import nodus

app = nodus.App("fn-warm")


@app.function(cpu=1, memory="1Gi", min_workers=1, max_workers=3, scaledown_window="2m", max_cost=1)
def ping() -> str:
    return "pong"
```

Terminal window

```console
$ nodus deploy examples/functions/warm-pool/app.py
```

```python
import nodus

ping = nodus.Function.from_name("fn-warm", "ping")
print(ping.remote())  # pong
```

A deploy updates the App in place. A Function whose code, image or resources changed rolls its workers once, after their in-flight calls finish. A Function you removed from the file is deleted, and the calls it still had queued end as `Failed` with the reason `FunctionDeleted`, so no caller waits for a call nobody will run.

## What Nodus creates

|Object|What it is|Look at it with|
|-|-|-|
|`App`|The Functions of one file; ephemeral for `run`, persistent for `deploy`|`nodus get apps`|
|`Function`|One decorated function or class, with its image, resources and `scaling`|`nodus get functions`|
|`FunctionCall`|One invocation, kept for seven days after it ends|`nodus get functioncalls -l nodus.dev/function=fn-warm-ping`|

A Function reports its state in `status.phase`, how many workers it has in `status.workers`, how many calls wait in `status.queue`, and the expected start time in `status.estimate`.

Terminal window

```console
$ nodus get function fn-warm-ping
NAME           PHASE     WORKERS   QUEUED   COST     AGE
fn-warm-ping   Running   1/3       0        $0.02    3m
```

## Stop, restart and delete

Terminal window

```console
$ nodus stop function/fn-warm-ping      # drain the workers; new calls wait in the queue
$ nodus start function/fn-warm-ping     # start workers again for the calls that waited
$ nodus restart function/fn-warm-ping   # replace every worker once, after in-flight calls finish
$ nodus delete app/fn-warm
```

A stopped Function keeps accepting calls and holds them in the queue, so work you submit while it is stopped runs once it starts. Workers also stop by themselves when your balance cannot cover another renewal; calls queue until you add credit and then run.

Note

A call that runs for longer than its `timeout` (five minutes by default, 24 hours at most) ends as `Failed` with the reason `DeadlineExceeded`, and its worker restarts, because the thread that ran it cannot be interrupted.


---

# What is billed while warm

> How Function workers are billed when they run calls, when they wait, and when they scale to zero.

Source: https://nodus-platform-site.pages.dev/docs/guides/functions/billing/
Build revision: 211ad9f836655b1c3a2668c4693e442471f28614

You pay for worker time, by the second, at the worker’s rate. A call does not add a charge of its own. A Function shows the worker’s rate before you run anything, in `status.estimate.rateUSDPerHour` and in `f.estimate(...)`.

## When a worker is billed

|State|Billed?|
|-|-|
|A worker starting: placing, pulling its image, running `enter` hooks|Yes, from the moment it is placed.|
|A worker running at least one call|Yes.|
|An idle worker inside `min_workers`|Yes, always. This is warm capacity.|
|An idle worker above `min_workers`, inside `scaledown_window`|Yes, until the window ends.|
|A worker that was released|No.|
|A Function with no workers|No.|

So a Function with `min_workers=0` costs nothing while nothing is calling it, and each burst costs its workers’ time plus their `scaledown_window`. A Function with `min_workers=1` costs one worker’s rate every hour, and its calls never wait for a start.

## Reading the bill

Every worker is billed under one Function and shows up as two lines of usage: the time the worker spent running calls and the time it spent waiting (`warm_idle_seconds`). Both appear under the Function in your usage records, so the cost of keeping workers warm is never mixed into the calls themselves.

Each call’s `status.costUSD` is the worker time attributed to that call: the worker’s rate divided by `target_concurrency`, for the whole seconds the call spent on a worker, at least one. A worker that runs four calls at once shows each of them a quarter of the rate. Call costs show where the busy time went. The bill is the workers’.

## Spending limits

A Function does not take a spend cap yet: setting `max_cost` is refused when you deploy, so no cap can be set and silently ignored. Your credit balance is the limit on what its workers spend, and an idle `min_workers` worker bills until you stop the Function.

When your balance cannot cover another renewal, the workers drain inside the reserve and the Function scales to zero. New calls queue and show `Funded=False` on the Function. Add credit and the Function starts workers again for the calls that waited. Nothing the calls were doing is lost: a call whose worker drained finishes first, and one that could not finish goes back to the queue.


---

# Call a Function

> Run one call, start calls without waiting, or map a Function over thousands of inputs, and read what comes back.

Source: https://nodus-platform-site.pages.dev/docs/guides/functions/calls/
Build revision: 211ad9f836655b1c3a2668c4693e442471f28614

A Function has three ways to start a call. Each call is a `FunctionCall` object, so you can look it up later, wait for it from another process or cancel it.

|Call|What it does|
|-|-|
|`f.remote(x)`|Runs one call and blocks until it returns its result.|
|`f.spawn(x)`|Starts one call and returns a handle: `handle.get(timeout=...)` waits, `handle.cancel()` stops it.|
|`f.map(xs)`|Starts one call per input and yields the results in input order.|

```python
with app.run():
    print(square.remote(7))                  # 49
    handle = square.spawn(8)
    print(handle.get(timeout=600))           # 64
    print(list(square.map(range(1000))))     # 1,000 results, in order
```

## Map over many inputs

`.map()` creates its calls in batches of up to 1,000 per request, and a batch is all or nothing: a request that fails leaves no calls behind, and sending it again does not create them twice. Results come back in input order whichever worker finishes first. `order_outputs=False` yields them as they finish, and `return_exceptions=True` yields a failed input’s exception instead of stopping the loop.

```python
results = list(process.map(files, order_outputs=False, return_exceptions=True))
```

Several callers can map into one Function at the same time. Workers take calls from each caller in turn, so a caller that submits 10 inputs is not stuck behind another caller’s 10,000.

## Arguments and results

Arguments and results are serialized with cloudpickle. Values up to 64 KiB travel inside the call. Larger ones are uploaded once, by content hash, and the call carries a reference, so a large argument costs nothing extra to send again. The caller and the worker image must run the same Python minor version, 3.10 to 3.13.

## Errors, retries and timeouts

An exception raised inside the Function is raised again in your process as its own type when that type can be imported there. Otherwise you get `nodus.errors.RemoteError`. Either way the remote traceback is attached.

```python
try:
    divide.remote(1, 0)
except ZeroDivisionError:
    ...
```

Two things can send a call back to the queue, and they are counted separately:

* **The function raised.** With `retries=nodus.Retries(max_retries=3)` the call is retried with exponential backoff, up to ten times. Without retries the exception ends the call.
* **The worker was lost.** A worker that disappears or is replaced does not count as a retry. Its calls go back to the queue and run on another worker, up to `recovery.maxAttempts` times (eight by default), then end as `Failed` with the reason `RecoveryLimitExceeded`.

A call can finish only once. A result that arrives from a worker that no longer holds the call is refused, so a call that was re-dispatched never ends with two results.

## Look a call up later

```python
call = nodus.FunctionCall.from_name("fn-warm-ping-bcdfghjklm")
print(call.get(timeout=60))
```

Terminal window

```console
$ nodus get functioncalls -l nodus.dev/map=<map id>
$ nodus get functioncall <name> -o yaml    # phase, result reference, retries, recoveries, costUSD
```

A call keeps its result for seven days after it ends. Each call reports `status.costUSD`, the worker time attributed to it; [the billing page](https://nodus-platform-site.pages.dev/docs/guides/functions/billing) explains how it relates to what you pay.


---

# Scaling and warm workers

> Set how many workers a Function keeps, how long an idle one stays, and how many calls each runs at once.

Source: https://nodus-platform-site.pages.dev/docs/guides/functions/scaling/
Build revision: 211ad9f836655b1c3a2668c4693e442471f28614

A Function’s `scaling` bounds its worker pool. Nodus sizes the pool from the calls that are waiting and running.

|Setting|Default|What it does|
|-|-|-|
|`min_workers`|0|Workers kept warm even when idle. [They are billed.](https://nodus-platform-site.pages.dev/docs/guides/functions/billing)|
|`max_workers`|10|The most workers the Function ever has, 1 to 1,000.|
|`scaledown_window`|`1m`|How long an idle worker above `min_workers` stays before it is released, up to 20 minutes.|
|`target_concurrency`|1|Calls one worker runs at once, 1 to 1,000. Use more than 1 for I/O-bound functions.|

```python
@app.function(cpu=2, memory="4Gi", min_workers=1, max_workers=20, scaledown_window="5m", target_concurrency=4)
def embed(text: str) -> list[float]:
    ...
```

A pool with `target_concurrency=4` holds one worker for every four calls that are waiting or running, and never more than `max_workers`. A worker leaves only after it has been idle for the whole `scaledown_window`, so a burst that comes back inside the window finds its workers still there.

## Cold and warm starts

A call that lands on a warm worker with a free slot starts at once. A call that needs a new worker waits for the worker to place, pull its image and start. Nodus shows both before you run anything:

```python
print(embed.estimate("hello"))   # expected cost, cold and warm start times, and the hold
```

Terminal window

```console
$ nodus get function embed -o jsonpath={.status.estimate}
```

`status.estimate.startup` holds the cold and warm bands, and `status.estimate.rateUSDPerHour` the worker’s rate. While workers are starting, the Function’s `Ready` condition says `WorkersStarting`. With no workers and nothing queued it says `ScaledToZero`, and with an image that is still building it says `ImagePending`.

## Classes: set up once per worker

A class keeps its state for the life of a worker. `@nodus.enter()` methods run once when the worker starts, before its first call. `@nodus.exit()` methods run once when the worker drains. Methods marked `@nodus.method()` get `.remote()`, `.map()` and `.spawn()`.

examples/python/classes/app.py

```python
"""A class whose model loads once per worker, then serves many calls.

Run it with `nodus run examples/python/classes/app.py`.
"""

import nodus

app = nodus.App("classes")


@app.cls(cpu=2, memory="4Gi", scaledown_window="5m", max_cost=1)
class Greeter:
    @nodus.enter()
    def load(self) -> None:
        # Runs once when a worker starts, before its first call: load weights or open connections here.
        self.greeting = "hello"

    @nodus.method()
    def greet(self, name: str) -> str:
        return f"{self.greeting}, {name}"

    @nodus.exit()
    def close(self) -> None:
        self.greeting = ""


@app.local_entrypoint()
def main() -> None:
    greeter = Greeter()
    print(greeter.greet.remote("Ada"))
    print(list(greeter.greet.map(["Grace", "Linus"])))
```

Load models and open connections in `enter`, so a call pays for them once per worker instead of once per call. A failing `exit` hook is logged and does not stop the others.

## Changing a Function

Editing `scaling` takes effect on the next reconcile without restarting any worker. Changing the code, image, Python version or resources rolls the workers once: new workers start, and the old ones finish their calls and leave. `nodus restart` rolls them without a change.

A Function whose image is still building keeps the workers it already has and keeps serving calls until the new image is ready.


---

# Choose GPUs and check availability

> See which accelerators are available right now, what they cost, and how to ask for exactly the hardware your run needs.

Source: https://nodus-platform-site.pages.dev/docs/guides/gpus/
Build revision: 211ad9f836655b1c3a2668c4693e442471f28614

You describe the hardware a run needs; Nodus finds capacity that fits and finishes it for the lowest expected cost. This guide shows how to see what is available, read an offering, and write a GPU request that says exactly what you mean.

## See what is available

Terminal window

```console
$ nodus get gpus
NAME                  TYPE       COUNT   VCPU   MEMORY   REGION    PRICE/H   AVAILABILITY   AGE
a100-sxm-80g-x1-any   A100-80G   1       30     200Gi    unknown   $1.27     Available      14s
h100-sxm-80g-x8-us    H100-SXM   8       208    1800Gi   us        $18.29    Available      14s
l4-24g-x1-eu-int      L4         1       8      32Gi     eu        $0.31     Limited        14s
```

`PRICE/H` is the from rate per machine-hour and `AGE` is how long ago the price and availability were confirmed.

Filter with flags; each one narrows the list:

Terminal window

```console
$ nodus get gpus --gpu H100 --count 8 --region us
$ nodus get gpus --interruptible
$ nodus get gpus --gpu L4 -o wide        # typical and list rates, startup time and interruption rate
```

The console’s GPU picker shows the same list.

## Read an offering

An offering is a class of capacity: accelerator × count × region class, and whether it is interruptible. Its name spells that out:

|Name|Meaning|
|-|-|
|`h100-sxm-80g-x8-us`|Eight H100 SXM 80 GB GPUs per machine, in the `us` region class|
|`l4-24g-x1-eu-int`|One L4 per machine in `eu`, interruptible|
|`a100-sxm-80g-x1-any`|Capacity with no guaranteed location (`any`)|
|`cpu-8c-32g-us`|An 8 vCPU, 32 GiB CPU machine in `us`|

Each offering shows:

* **From** and **typical** rates per hour. The from rate is that of the cheapest machine free now, the one a run starts on by default, or of the cheapest machine when none is free; the host shape shown is that machine’s. Your rate is shown in the estimate before launch and frozen for the run. No rate is ever above the list price.
* **Availability**: `Available`, `Limited` (a few machines, or a count that is not reported) or `Unavailable`.
* **Startup**: how long a machine usually takes to be ready (p50 and p90), from what Nodus has measured.
* **Interruption rate** for interruptible offerings: how often such capacity is taken back per hour.

Prices and availability refresh continuously. The public [pricing page](https://nodus-platform-site.pages.dev/pricing) shows the same from rates.

## Region classes

`placement.regions` takes region classes, not data centers: `us`, `ca`, `eu`, `uk`, `in`, `apac`, `me` and `latam`. Capacity whose location is not reported has the class `any` in its name and never satisfies a non-empty `placement.regions`, so a run that must stay in a region never lands there.

```yaml
spec:
  placement:
    regions: [eu, uk]
```

## Ask for the hardware you need

`resources.gpu.type` lists every accelerator you accept. Order does not matter: the scheduler picks among them by expected cost to finish.

|You write|Nodus may use|
|-|-|
|`H100`|Any H100 variant: SXM, PCIe or NVL|
|`[H100, H200]`|Any H100 or H200, never another family|
|`H100-SXM`|Only the H100 SXM|
|`H100!` or `exact: true`|Only the family’s primary variant (H100 SXM)|
|`minMemory: 80Gi` with no type|Any accelerator with at least 80 GiB per GPU|

Modal spellings work as written: `A100-40GB`, `A100-80GB`, `A10G`, `L40S`, `H100`, `H200`, `B200`, `T4`, `L4`.

```yaml
spec:
  resources:
    gpu: {type: [H100, H200], count: 8}
  placement:
    interruptible: Allow          # use interruptible capacity only when it is cheaper to finish
    maxRateUSDPerHour: "30.00"
```

`count` is per machine (1, 2, 4 or 8). For more GPUs than one machine holds, see [multi-node training](https://nodus-platform-site.pages.dev/docs/guides/multi-node/) (Beta).

## Use every GPU of one machine

`--gpu A6000:2` (or `count: 2`) gives your container both GPUs of one machine, numbered 0 and 1. `torchrun` starts one process per GPU without flags, because Nodus sets `PET_NPROC_PER_NODE` to the GPU count; a `--nproc-per-node` you pass yourself wins.

train.py

```python
"""A DDP smoke run on every GPU of one machine. torchrun starts one process per GPU (PET_NPROC_PER_NODE)."""
import os

import torch
import torch.distributed as dist
from torch.nn.parallel import DistributedDataParallel as DDP

dist.init_process_group("nccl")
local_rank = int(os.environ["LOCAL_RANK"])
device = torch.device("cuda", local_rank)
torch.cuda.set_device(device)

# Every process contributes 1, so the sum is the world size.
one = torch.ones(1, device=device)
dist.all_reduce(one)

model = DDP(torch.nn.Linear(16, 1).to(device), device_ids=[local_rank])
opt = torch.optim.SGD(model.parameters(), lr=0.1)
for step in range(20):
    x = torch.randn(32, 16, device=device)
    loss = (model(x) - x.sum(dim=1, keepdim=True)).pow(2).mean()
    opt.zero_grad()
    loss.backward()
    opt.step()

if dist.get_rank() == 0:
    print(f"world_size={int(one.item())} gpus={torch.cuda.device_count()} "
          f"device={torch.cuda.get_device_name(device)} loss={loss.item():.4f}")
dist.destroy_process_group()
```

run.sh

```sh
nodus run --name torchrun-multi-gpu --gpu A6000:2 --max-cost 1.00 -- torchrun train.py
```

The run prints `world_size=2 gpus=2` from the first process. The processes reach each other over NCCL on the machine itself; `/dev/shm` is sized for it (half the container’s memory), so you need no `--shm-size` or `--ipc=host`.

## Check before you launch

A dry run returns the estimate without starting anything:

Terminal window

```console
$ nodus apply -f job.yaml --dry-run=server -o estimate
```

It shows the offering Nodus would use, the expected cost p50 and p90, the startup time, the first hold and the minimum charge, and how long the estimate stays valid. When nothing fits, it says why, for example `CapacityUnavailable`, `NoListPrice`, `ExceedsRemainingBudget` or `MissesDeadline`, with a fix.

## When a run waits in Queued

A run with no matching capacity stays `Queued` and is placed as soon as capacity appears, until `placement.queueTimeout` (24 hours by default). `nodus describe` shows the offerings that were considered and why each was rejected. To start sooner, widen `resources.gpu.type` or `placement.regions`, allow interruptible capacity, or raise `placement.maxRateUSDPerHour`.

How Nodus chooses among offerings is explained in [placement and scheduling profiles](https://nodus-platform-site.pages.dev/docs/concepts/supply-placement/).


---

# Jobs

> Run a command to completion on a GPU or CPU, watch it, and download what it wrote.

Source: https://nodus-platform-site.pages.dev/docs/guides/jobs/
Build revision: 211ad9f836655b1c3a2668c4693e442471f28614

A Job runs your command until it finishes. You say what it needs (a GPU type, memory, a time limit, a spending cap) and Nodus runs it on capacity that fits, streams its logs, collects its outputs and charges only for the time it used.

## Sign in

Terminal window

```console
$ pip install nodus-compute
$ nodus login
```

`nodus login` opens the console in your browser and stores a key for this machine.

## Submit a Job

The fastest way is `nodus run`. It uploads the current directory, starts the command and follows it until it exits, passing the exit code through:

Terminal window

```console
$ nodus run --gpu L4 --image nodus/pytorch -- python hello.py
```

hello.py

```python
import json
import os

import torch

name = torch.cuda.get_device_name(0)
print(f"Hello from {name}")

# Anything written under /nodus/outputs is collected when the Job succeeds.
report = {"gpu": name, "cuda": torch.version.cuda, "index": os.environ["NODUS_INDEX"]}
with open(os.path.join(os.environ["NODUS_OUTPUTS_DIR"], "report.json"), "w") as f:
    json.dump(report, f)
```

Before anything runs, `nodus run` prints the estimate: the expected cost to completion and when it should start. Add `--max-cost 5` to stop the Job if it would spend more than $5, and `--timeout 2h` to limit its wall time.

To keep a Job in version control, write it as a manifest and apply it:

job.yaml

```yaml
apiVersion: nodus.dev/v1
kind: Job
metadata:
  name: hello
spec:
  image: nodus/pytorch
  command:
    - python
    - -c
    - |
      import json, os, torch
      name = torch.cuda.get_device_name(0)
      print(f"Hello from {name}")
      with open(os.path.join(os.environ["NODUS_OUTPUTS_DIR"], "report.json"), "w") as f:
          json.dump({"gpu": name, "index": os.environ["NODUS_INDEX"]}, f)
  resources:
    gpu: L4
  timeout: 10m
  maxCostUSD: "0.25"
  outputs:
    - name: report
      path: /nodus/outputs/report.json
```

Terminal window

```console
$ nodus apply -f job.yaml
job.nodus.dev/hello created
```

Anything under `/nodus/outputs` is collected when the Job succeeds. Declare the paths you want to download by name under `outputs`. `kubectl apply -f job.yaml` works too once your kubeconfig points at Nodus.

You do not have to name a GPU. Set `resources.gpu.minMemory` instead of `resources.gpu.type` and Nodus picks any accelerator with at least that much memory per GPU. If you also leave out `minMemory`, the annotations `nodus.dev/model` (for example `meta-llama/Llama-3.1-8B`) and `nodus.dev/dataset-bytes`, which you set under `metadata.annotations`, set it for you: 24Gi for 7B and 8B models, 32Gi for 13B and 14B, 48Gi for 34B and 40B, 80Gi for 70B and 72B, and 80Gi for a dataset over 1 TiB. `nodus apply -f job.yaml --dry-run=server -o yaml` shows the floor it picked.

## Watch it

In the console, open **Runs**, then select a run. **Logs** shows reported progress, stored outputs and the latest recovery checkpoint alongside the log stream. Distributed and indexed runs can be narrowed to a worker rank or index. The stream shows one worker at a time, starting with rank 0 and index 0. It follows that worker’s latest attempt; it is not a history of every retry. Connection failures show an error and **Retry logs**, while **Download loaded logs** saves the lines currently loaded in the browser. **Run activity** shows scheduling and lifecycle events separately from program output.

The command field preserves quoted arguments such as `python -c "print(1 + 1)"`. Shell pipelines and redirects need an explicit shell command, for example `bash -lc 'python prepare.py && python train.py'`.

Terminal window

```console
$ nodus get job/hello -w
$ nodus logs job/hello -f
$ nodus describe job/hello
```

`get` shows the phase, the GPU, the attempt, the cost so far and the age (`-o wide` adds the offering, its rate and progress). `describe` adds the conditions, the events and the reason for a failure with a suggested fix. A Job moves through these phases:

|Phase|What is happening|
|-|-|
|`Queued`|Waiting for capacity that fits and for a funded hold|
|`Provisioning`|Capacity is acquired; the image, source and inputs are being prepared|
|`Running`|Your command is running|
|`Recovering`|The capacity was lost; Nodus is moving the Job to new capacity|
|`Suspending`, `Suspended`|Saving state and releasing compute, then paused|
|`Cancelling`|Stopping and releasing compute|
|`Succeeded`, `Failed`, `Cancelled`|Finished; nothing is running or billed|

[Lifecycles](https://nodus-platform-site.pages.dev/docs/concepts/lifecycles/) lists every transition.

## Get the results

Open **Outputs** to download stored files and inspect their size and SHA-256. Distributed outputs identify their rank, and indexed outputs identify their index. A failed or cancelled run may have no stored outputs; its logs and activity explain what happened. **Details** includes attempts and the execution cost breakdown.

Terminal window

```console
$ nodus cp job/hello:outputs/report ./report.json
./report.json: 214 bytes, sha256 4f1c0e9a2b7d
```

`nodus cp` checks every download against the SHA-256 recorded when the Job finished. Outputs stay available until you delete the Job. [Outputs](https://nodus-platform-site.pages.dev/docs/guides/outputs/) covers directories, Indexed Jobs and loading results into a database.

## Run many indexes

An Indexed Job runs the same command `completions` times, at most `parallelism` at once. Each run sees its number in `NODUS_INDEX` (and in `JOB_COMPLETION_INDEX`, as on Kubernetes):

job.yaml

```yaml
apiVersion: nodus.dev/v1
kind: Job
metadata:
  name: indexed-outputs
spec:
  image: python:3.12-slim
  # Four indexes, two at a time. Each index sees its number in NODUS_INDEX.
  completions: 4
  parallelism: 2
  command:
    - python
    - -c
    - |
      import json, os
      i = int(os.environ["NODUS_INDEX"])
      rows = [{"index": i, "n": n, "square": n * n} for n in range(i * 100, (i + 1) * 100)]
      os.makedirs("/nodus/outputs/shards", exist_ok=True)
      with open(f"/nodus/outputs/shards/part-{i}.jsonl", "w") as f:
          f.writelines(json.dumps(r) + "\n" for r in rows)
      print(f"index {i}: wrote {len(rows)} rows")
  resources:
    cpu: "2"
    memory: 4Gi
  backoffLimit: 2
  timeout: 15m
  maxCostUSD: "0.10"
  outputs:
    - name: shards
      path: /nodus/outputs/shards
```

The Job succeeds when every index has succeeded. `backoffLimit` is how many failed runs the whole Job tolerates before it fails; a retried index starts fresh on new capacity. `status.completedIndexes` lists the finished indexes, such as `0-2,5`.

## Limit time and cost

|Field|Flag|What it does|
|-|-|-|
|`maxCostUSD`|`--max-cost`|The most the Job may spend. At the cap it saves its state and becomes `Suspended` with reason `MaxCostReached`; raise the cap to resume it. The cap can only be raised.|
|`timeout`|`--timeout`|Wall-clock limit counted from the first `Provisioning`; the Job fails with `DeadlineExceeded`. Time spent `Suspended` does not count. You can raise, lower or remove it on a running Job, and `status.timeoutTime` moves with it. `activeDeadlineSeconds` is accepted as an alias.|
|`completeByTime`|`--complete-by`|When you need the result. Nodus picks capacity that should finish in time.|
|`expectedDuration`|`--expected-duration`|Your estimate of the run time, used for the cost estimate before the Job has any history.|
|`placement.queueTimeout`||How long to wait for capacity before failing with `CapacityUnavailable`. The wait starts again when a suspended Job resumes.|

When credits run out or a budget is reached, a Job saves its state and becomes `Suspended` with reason `InsufficientCredits` or `BudgetExceeded`. It stops by `status.stopByTime`, the time its funds run out, even if the save is not finished, and resumes on its own once it is funded again.

## Suspend, resume and cancel

Terminal window

```console
$ nodus suspend job/train     # save state, release compute, stop billing
$ nodus resume job/train      # continue from the saved state
$ nodus cancel job/train      # stop for good
```

These set `spec.state` to `Suspended`, `Running` or `Cancelled`, so the same change works from a manifest. If a suspend cannot save state, the Job goes back to `Running` and its `Suspended` condition says `SuspendFailed`; nothing is lost. Nodus tries again ten minutes later, up to three tries in all (`status.suspendFailures` counts the failures), and then leaves the Job running. A Job with `recovery.continuity: Ephemeral` keeps no state, so suspending it restarts it from the beginning, and `nodus suspend` warns first. `nodus delete job/x` cancels a running Job before removing it.

## Recovery

Capacity can be lost while a Job runs, for example when interruptible capacity is reclaimed. What happens next depends on `recovery.continuity`:

|Continuity|After a loss|
|-|-|
|`Checkpointed` (default)|Resumes on new capacity from the last saved state in `NODUS_CHECKPOINT_DIR` (`/nodus/state`)|
|`Restartable`|Starts again from the beginning on new capacity|
|`Ephemeral`|Fails; nothing is kept|

Your program writes and loads its own state files in `NODUS_CHECKPOINT_DIR`; Nodus saves and restores that directory. Set `recovery.onInterruption: Fail` to fail instead of recovering. An index that uses up `recovery.maxAttempts` fails the Job with `RecoveryLimitExceeded`.

## When a Job fails

`status.reason` says why and `status.fix` says what to change. The common reasons:

|Reason|Meaning|
|-|-|
|`BackoffLimitExceeded`|Your command exited non-zero more times than `backoffLimit` allows|
|`OOMKilled`|The command ran out of memory; request more `resources.memory` or a larger GPU|
|`DeadlineExceeded`|`timeout` elapsed|
|`CapacityUnavailable`|No capacity fit the request within `placement.queueTimeout`|
|`ImagePullFailed`|The image does not exist or its registry refused the pull|
|`InvalidOutputPath`|A declared output was not written|
|`NoProgress`|Two recoveries in a row each ran less than `recovery.minProgressDuration` before losing capacity|
|`StorageUnavailable`|Nodus could not issue the storage credentials your command needs; run the Job again later|

## Clean up

Finished Jobs are deleted `ttlSecondsAfterFinished` seconds after they finish. Jobs from `nodus run` default to 30 days; pass `--keep` to keep one. Deleting a Job deletes its outputs.

## Multi-node Jobs

Beta

`spec.distributed` runs one Job across several nodes as a gang, with the rank, world size and rendezvous address in the environment. The Multi-node training guide covers launchers, networking and outputs from each rank.

## Environment

Every run sees these variables. Your own `env` cannot use the `NODUS_` prefix, except `NODUS_PARAM_`.

|Variable|Value|
|-|-|
|`NODUS_JOB`|The Job’s name|
|`NODUS_INDEX`, `JOB_COMPLETION_INDEX`|This run’s index, from 0|
|`NODUS_COMPLETIONS`|`spec.completions`|
|`NODUS_OUTPUTS_DIR`|`/nodus/outputs`|
|`NODUS_CHECKPOINT_DIR`|The first checkpoint path, `/nodus/state` by default|
|`NODUS_INPUT_<NAME>`|Where the input `<name>` is mounted, under `/nodus/inputs`|
|`NODUS_PARAM_<NAME>`|A Sweep parameter (see [Sweeps](https://nodus-platform-site.pages.dev/docs/guides/sweeps/))|

## Run code from GitHub

Install a GitHub Connection in the Job’s project and wait until it is Ready. Set `source.git.repo` to `owner/repository` or its GitHub HTTPS URL, and set `ref` to a branch, tag or commit. Nodus records the resolved commit and source archive in `status.source` before placing the Job. Retries and parallel indexes use the same files even when the branch moves.

```yaml
source:
  git:
    repo: your-org/your-repository
    ref: main
```

The files appear in `workingDir` before the command starts. Set `source.git.path` to use another checkout directory; this does not change the command’s working directory. The checkout contains repository files, without `.git` history, submodule contents or Git LFS downloads. GitHub credentials stay outside the guest. The source archive is limited to 100 MiB compressed, 512 MiB expanded and 10,000 entries. A missing Connection or rejected archive prevents execution.


---

# Sandboxes

> Create an isolated container, run commands in it with streaming output, read and write files, and stop it to stop paying for compute.

Source: https://nodus-platform-site.pages.dev/docs/guides/sandboxes/
Build revision: 211ad9f836655b1c3a2668c4693e442471f28614

A Sandbox is an isolated, long-running container for agents and untrusted code. You drive it with commands and file requests. When nothing uses it for a while it stops, keeps its `/workspace` directory and stops costing compute; the next command or file request starts it again. You pay per second while it holds compute and nothing while it is stopped.

Sandboxes run on CPU machines Nodus operates, each in its own isolated runtime with no network access unless you open it. [Sandbox isolation](https://nodus-platform-site.pages.dev/docs/concepts/sandboxes-isolation/) explains what keeps one Sandbox from another.

## Quick start

This manifest is the whole Sandbox. It asks for 1 vCPU and 2 GiB of memory, stops after 5 idle minutes and caps its spending at 1 USD:

sandbox.yaml

```yaml
apiVersion: nodus.dev/v1
kind: Sandbox
metadata:
  name: hello
spec:
  image: nodus/agent-tools # required: Python 3.12, Node 22 and git
  resources:
    cpu: "1"
    memory: 2Gi
  network:
    egress:
      policy: Deny # the default: no outbound traffic
  lifecycle:
    idleTimeout: 5m # stop after 5 minutes with no exec, file request or running process
    onIdle: Stop # keep /workspace and stop paying for compute; Delete removes the Sandbox instead
    maxLifetime: 2h # delete it 2 hours after creation whatever it is doing
  maxCostUSD: "1.00"
```

Create it, run a command, copy a file in and read it back:

Terminal window

```console
$ nodus apply -f sandbox.yaml
sandbox/hello created
$ nodus exec sb/hello -- echo hello from a sandbox
hello from a sandbox
$ nodus cp ./notes.txt sb/hello:/workspace/notes.txt
$ nodus exec sb/hello -- cat /workspace/notes.txt
written from my laptop
$ nodus stop sb/hello
$ nodus exec sb/hello -- cat /workspace/notes.txt    # starts it again; /workspace is kept
written from my laptop
$ nodus delete sb/hello
```

From Python, the same steps are in the [Python Sandboxes guide](https://nodus-platform-site.pages.dev/docs/guides/python/sandboxes/).

## Create a Sandbox

`nodus apply -f sandbox.yaml`, `nodus create sandbox <name>` and the SDK all create the same object. `image` is required on the API; the CLI and the SDK fill in `nodus/agent-tools` (Python 3.12, Node 22 and git) when you leave it out.

|Field|Default|Accepted values|
|-|-|-|
|`spec.image`|— (required)|Any image reference, or a Nodus catalog image such as `nodus/agent-tools`|
|`spec.resources.cpu`|`1`|0.25 to 64 vCPU|
|`spec.resources.memory`|`1Gi`|128 MiB to 256 GiB|
|`spec.resources.disk`|`10Gi`|1 GiB to 200 GiB|
|`spec.network.egress.policy`|`Deny`|`Deny` (no outbound traffic) or `Open` (public internet)|
|`spec.lifecycle.idleTimeout`|`5m`|`0s` (never idle) or 1 minute to 24 hours|
|`spec.lifecycle.onIdle`|`Stop`|`Stop` or `Delete`|
|`spec.lifecycle.maxLifetime`|`24h`|1 minute to 720 hours|
|`spec.continuity.mode`|`Snapshotted`|`Snapshotted` (keep `/workspace` across stops) or `Ephemeral` (start empty)|
|`spec.workingDir`|`/workspace`|An absolute path|
|`spec.maxCostUSD`|none|A USD amount such as `"5.00"`; the Sandbox stops when it has cost this much|
|`spec.state`|`Running`|`Running` or `Stopped`|

`nodus get sb/hello` shows `PHASE` (`Pending`, `Starting`, `Running`, `Stopping`, `Stopped`, `Recovering`, `Terminating`, `Failed`), `ACTIVITY` (`Busy` or `Idle`), the `CPU`, `MEMORY` and `GPU` you asked for, and `COST` so far.

### Create by name reconnects

Creating a Sandbox whose name already exists in the project does not make a second one:

* **Same spec.** You get the existing Sandbox back, and if it is stopped it starts again. This is how an agent reconnects to its Sandbox after a restart: it creates the same manifest every time.
* **A `maxCostUSD` that is added or raised.** The new budget is applied; nothing else changes. A budget can only be raised: a lower one, or none when the Sandbox has one, counts as a different spec.
* **Any other change.** The request fails with `409 AlreadyExists` and lists the fields that differ in `details.diff` and `details.causes`. Delete the Sandbox and create it again to change its image, shape or lifecycle.

GPU Sandboxes are not available yet: a spec with `resources.gpu` is refused.

## Run commands

`nodus exec` runs one command in the running Sandbox and streams its standard output and standard error back as the command writes them. The exit code of `nodus exec` is the command’s. Every command is recorded as a Process of the Sandbox, whatever started it (the CLI, an agent or the MCP tools): its name, such as `sb-hello-12`, comes back with the result, and `nodus get processes` lists it with how it ended.

Terminal window

```console
$ nodus exec sb/hello -- python -c 'print(6 * 7)'
42
$ nodus exec sb/hello -- sh -c 'pip list 2>/dev/null | head -3; exit 3'; echo "exit $?"
Package    Version
---------- -------
pip        24.2
exit 3
```

A command runs in `spec.workingDir` (`/workspace` unless you set it) with the Sandbox’s environment, as a non-root user. Pass a single string to `sh -c` to use pipes and redirection. Every command counts as activity, so a Sandbox never stops as idle while a command is running.

### Open a shell

`nodus shell` is `nodus create sandbox` and `nodus exec -it` in one step: it opens an interactive terminal in a Sandbox and creates the Sandbox first when the name is new. It takes the same flags as `nodus create sandbox`, and the image is `nodus/agent-tools` unless you pass `--image`.

Terminal window

```console
$ nodus shell sandbox/dev --cpu 2 --memory 4Gi
sandbox/dev created
sandboxes/dev is starting; waiting for it
```

You are now in a terminal in the Sandbox, in `/workspace`. When you exit, `nodus shell sandbox/dev` opens it again with `/workspace` as you left it.

The Sandbox stays after you exit, and stops by itself when idle. Pass `--rm` to delete a Sandbox this command created when the shell exits; `--rm` refuses a Sandbox that already exists, and so does a flag such as `--cpu`, because neither changes a Sandbox that is there. `nodus shell` with no name gives the shell a Sandbox of its own and deletes it on exit unless you pass `--keep`. Put a command after `--` to run it instead of `bash`: `nodus shell sandbox/dev -- zsh`. With input that is not a terminal, the shell runs without one, so `echo 'make test' | nodus shell sandbox/dev` works in a script. The exit code is the shell’s.

### Timeouts

Each command has its own timeout: 10 minutes unless you set one, and at most 24 hours. When the timeout passes, the process is killed with `SIGKILL` and the command ends with the reason `DeadlineExceeded`. The Sandbox keeps running. Set the timeout per command, for example `sb.exec("make", "test", timeout="2m")` in Python or `"timeout": "2m"` in the API request.

How a command ended is in its result:

|Reason|Meaning|
|-|-|
|none|Exit code 0|
|`NonZeroExit`|The command exited with another code|
|`Signaled`|A signal ended it; the exit code is 128 plus the signal number|
|`DeadlineExceeded`|Its timeout passed and it was killed|
|`StartFailed`|It could not start, for example because the program does not exist|
|`NodeLost`|The machine running the Sandbox was lost while the command ran|
|`OutputLimitExceeded`|It wrote more output than a Process may keep|
|`ParentStopped`|The Sandbox stopped or was deleted while the command ran|
|`Cancelled`|You cancelled it|

A command line holds at most 1,024 arguments and 64 KiB in total, plus up to 256 extra environment variables. A Sandbox runs at most 1,024 commands at once; the next one fails with `429 TooManyRequests` until one ends.

## Files

`nodus cp` copies files in both directions, and works on every image because the file service is part of the Sandbox runtime, not of the image:

Terminal window

```console
$ nodus cp ./notes.txt sb/hello:/workspace/notes.txt     # laptop to Sandbox
$ nodus cp sb/hello:/workspace/result.json ./result.json  # Sandbox to laptop
```

The same operations are one request each on `/apis/nodus.dev/v1/namespaces/{project}/sandboxes/{name}/files`:

|Request|Does|
|-|-|
|`GET …/files?path=/workspace/a.txt`|Returns the file’s bytes|
|`PUT …/files?path=/workspace/a.txt`|Replaces the file with the request body, creating parent directories|
|`DELETE …/files?path=/workspace/a.txt`|Deletes a file or an empty directory|

A write replaces the file atomically: a reader sees the old content or the new content, never half of it. A file is at most 64 MiB per request; a larger one fails with `413 FileTooLarge`. Paths are absolute, and a path that does not exist answers `404 NotFound`.

### Writing without overwriting someone else’s change

When two writers share a file, send the digest of the content you started from as `expectedSHA256`:

Terminal window

```console
$ curl -X PUT --data-binary @plan.md "$API/sandboxes/hello/files?path=/workspace/plan.md&expectedSHA256=$(sha256sum old-plan.md | cut -d' ' -f1)"
```

The write happens only if the file still has that digest. If it changed, nothing is written and the request fails with `412 SHA256Mismatch`: read the file again, merge, and retry. The digest is 64 lowercase hex digits. The digest of empty content (`e3b0c442…b855`) matches a file that does not exist yet, so it means “create only”.

## Stop, start and wake

A Sandbox is `Running` while it holds compute and `Stopped` while it does not. A stop saves `/workspace` first, so a later start gives it back, as long as `spec.continuity.mode` is `Snapshotted`, the default.

Terminal window

```console
$ nodus stop sb/hello        # saves /workspace, releases compute, ends the charge
$ nodus start sb/hello       # starts it again with /workspace restored
$ nodus get sb/hello
NAME    PHASE     ACTIVITY   CPU   MEMORY   GPU   COST      AGE
hello   Stopped   Idle       1     1Gi      -     $0.0042   12m
```

`nodus stop` sets `spec.state: Stopped` and the Sandbox stays stopped until you start it. Every other stop is made by the platform, shows its reason in `status.stopReason`, and ends the next time something needs the Sandbox:

|`stopReason`|Why it stopped|
|-|-|
|`User`|You ran `nodus stop`|
|`Idle`|Nothing used it for `idleTimeout`|
|`InsufficientCredits`|Your balance could not fund the next period|
|`BudgetExceeded`|A Budget on the project reached its limit|
|`MaxCostReached`|The Sandbox cost as much as `spec.maxCostUSD`|

A stopped Sandbox starts again on any of these wake triggers:

* a command (`nodus exec`, or an exec request),
* a file request,
* `nodus start`,
* creating a Sandbox with the same name and an identical spec.

Credits and budgets still apply to a wake. A Sandbox stopped for `InsufficientCredits`, `BudgetExceeded` or `MaxCostReached` stays stopped until the balance, the Budget or `maxCostUSD` allows the next period of compute. Until then a command, a file request and `nodus start` answer `402` instead of `SandboxStarting`: `InsufficientCredits` states the hold the start needs and what you have, and `BudgetExceeded` the Budget or `maxCostUSD` that is full. Add credits, raise the Budget or raise `maxCostUSD`, then send a trigger again. A create by name records its wake too, but returns the Sandbox as it is: still stopped, with `Funded` false in its conditions.

### `SandboxStarting` and `Retry-After`

A command or file request that finds the Sandbox stopped starts it and answers `503 SandboxStarting` with a `Retry-After` header (2 seconds), because the container is not ready yet. The CLI and the SDK wait and retry for you, so `nodus exec` on a stopped Sandbox just takes a few seconds longer. If you call the API directly, retry after the time in `Retry-After`; each retry is also a wake trigger, so it does no harm to send them early.

Other reasons a request can fail while the Sandbox is not usable:

|Status and reason|Meaning|
|-|-|
|`409 SandboxNotRunning`|The Sandbox is being deleted|
|`409 SandboxFailed`|The Sandbox failed to start or lost its machine; `status.reason` says why|
|`402 InsufficientCredits`|The Sandbox is stopped for lack of credits and cannot start yet|
|`402 BudgetExceeded`|The Sandbox is stopped because a Budget or its `maxCostUSD` is full and cannot start yet|
|`404 NotFound`|There is no Sandbox with that name|

Each code has a page in the [error reference](https://nodus-platform-site.pages.dev/docs/reference/errors/).

## Idle, `onIdle` and `maxLifetime`

A Sandbox is **idle** when no command is running, no file request arrived for `idleTimeout` and no terminal had input in that time. A running process, including one that prints nothing for hours, counts as busy; a terminal that is open but silent does not count as activity.

When the idle time passes, `onIdle` decides what happens:

* `Stop` (the default) saves `/workspace`, releases compute and sets `stopReason: Idle`. The next wake trigger starts it again.
* `Delete` removes the Sandbox with its files. Use it for one-off work that leaves nothing worth keeping.

`idleTimeout: 0s` turns idle detection off, so the Sandbox runs until you stop it, until `maxLifetime` or until a budget ends it.

`maxLifetime` is a hard deadline counted from creation, 24 hours by default. When it passes, the Sandbox is deleted whatever it is doing, running or stopped. It is the backstop that keeps a forgotten Sandbox from living forever. Create a new Sandbox, or one with a longer `maxLifetime`, to carry on.

## Network access

A Sandbox has no outbound network access unless you open it. `spec.network.egress.policy` is one of:

* `Deny` (the default): nothing leaves the Sandbox, so code in it cannot download or upload anything.
* `Open`: the Sandbox can reach the public internet. It still cannot reach other Sandboxes, the machine it runs on or Nodus’s own services.

A Sandbox accepts no inbound connections; you reach it only through commands and file requests.

## What a Sandbox costs

A Sandbox is billed per second, from the moment its machine slot is reserved until it is released, at the rate of its shape. The rate is the Nodus machine price for the CPU, memory and disk you asked for, and it already includes Nodus’s 12.5 % ([how pricing works](https://nodus-platform-site.pages.dev/docs/concepts/pricing/)). Four things to know:

* **Starting and stopping are billed.** The seconds spent starting (or restoring `/workspace`) and the seconds spent saving a stop are part of the reservation, shown as their own lines on the bill.
* **Stopped time is free.** A stopped Sandbox holds no compute, so no compute cost accrues. The saved `/workspace` counts as storage, which is metered separately from compute.
* **The rate is frozen per start.** The rate a start shows is the rate that start pays, even if prices change while it runs.
* **Shared capacity has limits.** New shared capacity starts for you at most four times an hour, and none starts for 30 minutes after capacity started for you goes unused; until your organization’s first purchase it starts only for Sandboxes of up to 3 vCPU and 11 GiB. Otherwise a Sandbox runs on a machine of its own, at that machine’s rate.

`nodus get sb/hello` shows the running total in the `COST` column. Set `spec.maxCostUSD` to cap one Sandbox: it stops when it has cost that much, and you can raise the cap with another create by name. Without a cap, your credit balance and any Budget on the project are the limits.

## Next steps

* [Sandbox isolation](https://nodus-platform-site.pages.dev/docs/concepts/sandboxes-isolation/) for what separates your Sandbox from the machine and from other Sandboxes.
* [Sandboxes from Python](https://nodus-platform-site.pages.dev/docs/guides/python/sandboxes/) for the SDK, including terminals and previews.


---

# Workspaces

> A GPU or CPU development machine with a saved home directory, reached with SSH, VS Code or JupyterLab.

Source: https://nodus-platform-site.pages.dev/docs/guides/workspaces/
Build revision: 211ad9f836655b1c3a2668c4693e442471f28614

A Workspace is your own development machine: a GPU or CPU machine with VS Code in the browser, JupyterLab and SSH. Its home directory, `/home/nodus`, lives on a Volume and is saved every time the Workspace stops, so you can stop it at the end of the day and pick up where you left off. While it is stopped you pay only for the saved files.

You need the `nodus` CLI to connect over SSH ([install it](https://nodus-platform-site.pages.dev/docs/getting-started/) and run `nodus login`). Browser VS Code and JupyterLab need only the console.

## Create a Workspace

Terminal window

```console
$ nodus create workspace lab --gpu H100
$ nodus get workspaces
NAME   PHASE     COMPUTE    HOME       STOP-REASON   COST    AGE
lab    Running   1 × H100   lab-home                 $0.09   3m
```

`--gpu H100:2` asks for two GPUs; `--cpu 8 --memory 32Gi` without `--gpu` creates a CPU Workspace. The home Volume is `lab-home`, created on the first start if it does not exist; set `spec.volume` to use another one. The default image is `nodus/workspace-pytorch-cuda` (PyTorch and CUDA) on NVIDIA GPUs, `nodus/workspace-rocm` on AMD GPUs and `nodus/workspace-cpu` on CPU. Every catalog image has VS Code, JupyterLab, Python and the user `nodus`.

The same Workspace as a manifest, applied with `nodus apply -f workspace.yaml`:

workspace.yaml

```yaml
apiVersion: nodus.dev/v1
kind: Workspace
metadata:
  name: ws-home
spec:
  image: nodus/workspace-cpu
  resources:
    cpu: "2"
    memory: 8Gi
  volume: ws-home-home
  idleTimeout: 30m
```

In the console, **Workspaces › New workspace** shows the GPU picker, the estimate and the same manifest as YAML, CLI and Python before you create it.

## Connect

### SSH

Terminal window

```console
$ nodus login
$ nodus ssh workspace/lab
nodus@lab:~$ nvidia-smi
$ nodus ssh workspace/lab -- python train.py --epochs 1
```

Sign in once with `nodus login` on each computer using an up-to-date CLI and OpenSSH 8.5 or newer. Login creates a dedicated local SSH key for your account, registers only its public key, and configures every existing and future CPU or GPU Workspace in the organizations you selected. New Workspaces need no separate setup. Your private key stays on your computer with owner-only permissions; sign-in does not start or connect to a Workspace.

Each teammate signs in to their own account. Organization roles and project access determine which Workspaces they can enter. Remove a device’s key under **Account › SSH keys** to revoke it. Removing a teammate’s access also prevents their keys from opening that organization’s Workspaces.

`nodus ssh` connects through the Nodus gateway. No port is open on the machine. OpenSSH verifies the Workspace’s host key against the authenticated API on every connection, with strict host-key checking. A missing or revoked identity fails closed; it does not trust a host key on first use.

### VS Code Remote-SSH, Cursor and JetBrains Gateway

After `nodus login`, open **Desktop VS Code** or **Cursor** in the console, or select the Workspace host in your editor’s Remote-SSH connection dialog. The same account key works for all your authorized Workspaces, including ones created later. Signing in to the website alone cannot configure SSH on your computer.

The `.nodus` host is an SSH alias, not a public DNS address. Login records the CLI executable and saved context, so editors do not depend on your terminal’s environment. Plain `ssh`, `scp` and `rsync` use the same connection. `nodus ssh workspace/lab --config` prints the rule; `nodus ssh workspace/lab --setup` is an optional repair command. Neither command starts or connects to a Workspace.

### Browser VS Code and JupyterLab

Open the Workspace in the console and choose **Browser VS Code** or **JupyterLab**. Both open on a private address that only members of your organization can reach after signing in.

## Stop, start and what is saved

Terminal window

```console
$ nodus stop workspace/lab
$ nodus start workspace/lab
```

When a Workspace stops, everything in `/home/nodus` is saved to its home Volume and the machine is released. Everything else is temporary, including the `/workspace` directory (the container’s working directory, which has nothing to do with the Workspace kind) and anything installed outside your home. Install Python packages with `pip install --user` or in a virtual environment under `/home/nodus` to keep them.

The next start restores the saved home on a fresh machine. Use `ephemeral: true` for a Workspace with no home Volume, whose files are deleted on stop.

A Workspace also stops by itself:

|Stop reason|When|How it starts again|
|-|-|-|
|`Idle`|No SSH session, editor traffic or process for `idleTimeout` (default `1h`)|Connect, open a tool, or `nodus start`|
|`Schedule`|At `schedule.stopAt`|The next `schedule.readyBy`, connecting, or `nodus start`|
|`InsufficientCredits`, `BudgetExceeded`, `MaxCostReached`|Money ran out or a limit was reached|Add credits or raise the limit, then connect or `nodus start`|
|`User`|You ran `nodus stop` or chose **Stop**|Only `nodus start` or **Start**|

Connecting to a Workspace that stopped by itself starts it; `nodus ssh` waits and connects when it is ready. The console asks before a connect starts a stopped Workspace and shows the rate it will bill at.

### Schedules

```yaml
spec:
  schedule:
    readyBy: "2026-10-01T09:00:00-07:00"
    stopAt: "2026-10-01T19:00:00-07:00"
```

The Workspace starts 15 minutes before `readyBy` so it is ready on time, and stops at `stopAt`. It records a `ScheduledStart` event, or `ScheduleMissed` when it could not start in time.

## Run a job on your Workspace files

Terminal window

```console
$ nodus run --from workspace/lab -- python train.py
```

The Job runs on a copy of the last saved home, labelled `nodus.dev/workspace=lab`, so the Workspace’s **Jobs from this workspace** list shows it.

## Sessions and billing

Each period a Workspace runs is a session, recorded as an Attempt with its receipt:

Terminal window

```console
$ nodus get attempts -l nodus.dev/workspace=lab
```

* A GPU Workspace bills at the rate shown when it starts, from when its machine is acquired until the machine is released. A session that stops for any reason closes its charge.
* A CPU Workspace bills at the listed CPU rates while it runs.
* Saved files in the home Volume bill as storage after your organization’s included storage.

`maxCostUSD` stops and saves the Workspace when its cost reaches the limit; raising the limit lets it start again.

## Shared access

A Workspace belongs to its project. Every member of the organization with access to the project can see it, start it and connect with their own account and SSH key; `spec.sshKeys` limits SSH to the keys you name.

## When a Workspace fails

A Workspace in the `Failed` phase does not restart. Delete it and create it again with the same name (or the same `spec.volume`): the home Volume keeps your files.

Terminal window

```console
$ nodus delete workspace/lab
$ nodus create workspace lab --gpu H100
```

To delete the saved files too, delete the home Volume (`nodus delete volume/lab-home`), or empty it with `nodus volume clear lab-home`.
