Skip to content

Sandboxes

View Markdown

A Sandbox is an isolated, long-running container for agents and untrusted code. You drive it with commands and file requests. When nothing uses it for a while it stops, keeps its /workspace directory and stops costing compute; the next command or file request starts it again. You pay per second while it holds compute and nothing while it is stopped.

Sandboxes run on CPU machines Nodus operates, each in its own isolated runtime with no network access unless you open it. Sandbox isolation explains what keeps one Sandbox from another.

This manifest is the whole Sandbox. It asks for 1 vCPU and 2 GiB of memory, stops after 5 idle minutes and caps its spending at 1 USD:

sandbox.yaml
apiVersion: nodus.dev/v1
kind: Sandbox
metadata:
name: hello
spec:
image: nodus/agent-tools # required: Python 3.12, Node 22 and git
resources:
cpu: "1"
memory: 2Gi
network:
egress:
policy: Deny # the default: no outbound traffic
lifecycle:
idleTimeout: 5m # stop after 5 minutes with no exec, file request or running process
onIdle: Stop # keep /workspace and stop paying for compute; Delete removes the Sandbox instead
maxLifetime: 2h # delete it 2 hours after creation whatever it is doing
maxCostUSD: "1.00"

Create it, run a command, copy a file in and read it back:

Terminal window
$ nodus apply -f sandbox.yaml
sandbox/hello created
$ nodus exec sb/hello -- echo hello from a sandbox
hello from a sandbox
$ nodus cp ./notes.txt sb/hello:/workspace/notes.txt
$ nodus exec sb/hello -- cat /workspace/notes.txt
written from my laptop
$ nodus stop sb/hello
$ nodus exec sb/hello -- cat /workspace/notes.txt # starts it again; /workspace is kept
written from my laptop
$ nodus delete sb/hello

From Python, the same steps are in the Python Sandboxes guide.

nodus apply -f sandbox.yaml, nodus create sandbox <name> and the SDK all create the same object. image is required on the API; the CLI and the SDK fill in nodus/agent-tools (Python 3.12, Node 22 and git) when you leave it out.

Field Default Accepted values
spec.image — (required) Any image reference, or a Nodus catalog image such as nodus/agent-tools
spec.resources.cpu 1 0.25 to 64 vCPU
spec.resources.memory 1Gi 128 MiB to 256 GiB
spec.resources.disk 10Gi 1 GiB to 200 GiB
spec.network.egress.policy Deny Deny (no outbound traffic) or Open (public internet)
spec.lifecycle.idleTimeout 5m 0s (never idle) or 1 minute to 24 hours
spec.lifecycle.onIdle Stop Stop or Delete
spec.lifecycle.maxLifetime 24h 1 minute to 720 hours
spec.continuity.mode Snapshotted Snapshotted (keep /workspace across stops) or Ephemeral (start empty)
spec.workingDir /workspace An absolute path
spec.maxCostUSD none A USD amount such as "5.00"; the Sandbox stops when it has cost this much
spec.state Running Running or Stopped

nodus get sb/hello shows PHASE (Pending, Starting, Running, Stopping, Stopped, Recovering, Terminating, Failed), ACTIVITY (Busy or Idle), the CPU, MEMORY and GPU you asked for, and COST so far.

Creating a Sandbox whose name already exists in the project does not make a second one:

  • Same spec. You get the existing Sandbox back, and if it is stopped it starts again. This is how an agent reconnects to its Sandbox after a restart: it creates the same manifest every time.
  • A maxCostUSD that is added or raised. The new budget is applied; nothing else changes. A budget can only be raised: a lower one, or none when the Sandbox has one, counts as a different spec.
  • Any other change. The request fails with 409 AlreadyExists and lists the fields that differ in details.diff and details.causes. Delete the Sandbox and create it again to change its image, shape or lifecycle.

GPU Sandboxes are not available yet: a spec with resources.gpu is refused.

nodus exec runs one command in the running Sandbox and streams its standard output and standard error back as the command writes them. The exit code of nodus exec is the command’s. Every command is recorded as a Process of the Sandbox, whatever started it (the CLI, an agent or the MCP tools): its name, such as sb-hello-12, comes back with the result, and nodus get processes lists it with how it ended.

Terminal window
$ nodus exec sb/hello -- python -c 'print(6 * 7)'
42
$ nodus exec sb/hello -- sh -c 'pip list 2>/dev/null | head -3; exit 3'; echo "exit $?"
Package Version
---------- -------
pip 24.2
exit 3

A command runs in spec.workingDir (/workspace unless you set it) with the Sandbox’s environment, as a non-root user. Pass a single string to sh -c to use pipes and redirection. Every command counts as activity, so a Sandbox never stops as idle while a command is running.

nodus shell is nodus create sandbox and nodus exec -it in one step: it opens an interactive terminal in a Sandbox and creates the Sandbox first when the name is new. It takes the same flags as nodus create sandbox, and the image is nodus/agent-tools unless you pass --image.

Terminal window
$ nodus shell sandbox/dev --cpu 2 --memory 4Gi
sandbox/dev created
sandboxes/dev is starting; waiting for it

You are now in a terminal in the Sandbox, in /workspace. When you exit, nodus shell sandbox/dev opens it again with /workspace as you left it.

The Sandbox stays after you exit, and stops by itself when idle. Pass --rm to delete a Sandbox this command created when the shell exits; --rm refuses a Sandbox that already exists, and so does a flag such as --cpu, because neither changes a Sandbox that is there. nodus shell with no name gives the shell a Sandbox of its own and deletes it on exit unless you pass --keep. Put a command after -- to run it instead of bash: nodus shell sandbox/dev -- zsh. With input that is not a terminal, the shell runs without one, so echo 'make test' | nodus shell sandbox/dev works in a script. The exit code is the shell’s.

Each command has its own timeout: 10 minutes unless you set one, and at most 24 hours. When the timeout passes, the process is killed with SIGKILL and the command ends with the reason DeadlineExceeded. The Sandbox keeps running. Set the timeout per command, for example sb.exec("make", "test", timeout="2m") in Python or "timeout": "2m" in the API request.

How a command ended is in its result:

Reason Meaning
none Exit code 0
NonZeroExit The command exited with another code
Signaled A signal ended it; the exit code is 128 plus the signal number
DeadlineExceeded Its timeout passed and it was killed
StartFailed It could not start, for example because the program does not exist
NodeLost The machine running the Sandbox was lost while the command ran
OutputLimitExceeded It wrote more output than a Process may keep
ParentStopped The Sandbox stopped or was deleted while the command ran
Cancelled You cancelled it

A command line holds at most 1,024 arguments and 64 KiB in total, plus up to 256 extra environment variables. A Sandbox runs at most 1,024 commands at once; the next one fails with 429 TooManyRequests until one ends.

nodus cp copies files in both directions, and works on every image because the file service is part of the Sandbox runtime, not of the image:

Terminal window
$ nodus cp ./notes.txt sb/hello:/workspace/notes.txt # laptop to Sandbox
$ nodus cp sb/hello:/workspace/result.json ./result.json # Sandbox to laptop

The same operations are one request each on /apis/nodus.dev/v1/namespaces/{project}/sandboxes/{name}/files:

Request Does
GET …/files?path=/workspace/a.txt Returns the file’s bytes
PUT …/files?path=/workspace/a.txt Replaces the file with the request body, creating parent directories
DELETE …/files?path=/workspace/a.txt Deletes a file or an empty directory

A write replaces the file atomically: a reader sees the old content or the new content, never half of it. A file is at most 64 MiB per request; a larger one fails with 413 FileTooLarge. Paths are absolute, and a path that does not exist answers 404 NotFound.

Writing without overwriting someone else’s change

Section titled “Writing without overwriting someone else’s change”

When two writers share a file, send the digest of the content you started from as expectedSHA256:

Terminal window
$ curl -X PUT --data-binary @plan.md "$API/sandboxes/hello/files?path=/workspace/plan.md&expectedSHA256=$(sha256sum old-plan.md | cut -d' ' -f1)"

The write happens only if the file still has that digest. If it changed, nothing is written and the request fails with 412 SHA256Mismatch: read the file again, merge, and retry. The digest is 64 lowercase hex digits. The digest of empty content (e3b0c442…b855) matches a file that does not exist yet, so it means “create only”.

A Sandbox is Running while it holds compute and Stopped while it does not. A stop saves /workspace first, so a later start gives it back, as long as spec.continuity.mode is Snapshotted, the default.

Terminal window
$ nodus stop sb/hello # saves /workspace, releases compute, ends the charge
$ nodus start sb/hello # starts it again with /workspace restored
$ nodus get sb/hello
NAME PHASE ACTIVITY CPU MEMORY GPU COST AGE
hello Stopped Idle 1 1Gi - $0.0042 12m

nodus stop sets spec.state: Stopped and the Sandbox stays stopped until you start it. Every other stop is made by the platform, shows its reason in status.stopReason, and ends the next time something needs the Sandbox:

stopReason Why it stopped
User You ran nodus stop
Idle Nothing used it for idleTimeout
InsufficientCredits Your balance could not fund the next period
BudgetExceeded A Budget on the project reached its limit
MaxCostReached The Sandbox cost as much as spec.maxCostUSD

A stopped Sandbox starts again on any of these wake triggers:

  • a command (nodus exec, or an exec request),
  • a file request,
  • nodus start,
  • creating a Sandbox with the same name and an identical spec.

Credits and budgets still apply to a wake. A Sandbox stopped for InsufficientCredits, BudgetExceeded or MaxCostReached stays stopped until the balance, the Budget or maxCostUSD allows the next period of compute. Until then a command, a file request and nodus start answer 402 instead of SandboxStarting: InsufficientCredits states the hold the start needs and what you have, and BudgetExceeded the Budget or maxCostUSD that is full. Add credits, raise the Budget or raise maxCostUSD, then send a trigger again. A create by name records its wake too, but returns the Sandbox as it is: still stopped, with Funded false in its conditions.

A command or file request that finds the Sandbox stopped starts it and answers 503 SandboxStarting with a Retry-After header (2 seconds), because the container is not ready yet. The CLI and the SDK wait and retry for you, so nodus exec on a stopped Sandbox just takes a few seconds longer. If you call the API directly, retry after the time in Retry-After; each retry is also a wake trigger, so it does no harm to send them early.

Other reasons a request can fail while the Sandbox is not usable:

Status and reason Meaning
409 SandboxNotRunning The Sandbox is being deleted
409 SandboxFailed The Sandbox failed to start or lost its machine; status.reason says why
402 InsufficientCredits The Sandbox is stopped for lack of credits and cannot start yet
402 BudgetExceeded The Sandbox is stopped because a Budget or its maxCostUSD is full and cannot start yet
404 NotFound There is no Sandbox with that name

Each code has a page in the error reference.

A Sandbox is idle when no command is running, no file request arrived for idleTimeout and no terminal had input in that time. A running process, including one that prints nothing for hours, counts as busy; a terminal that is open but silent does not count as activity.

When the idle time passes, onIdle decides what happens:

  • Stop (the default) saves /workspace, releases compute and sets stopReason: Idle. The next wake trigger starts it again.
  • Delete removes the Sandbox with its files. Use it for one-off work that leaves nothing worth keeping.

idleTimeout: 0s turns idle detection off, so the Sandbox runs until you stop it, until maxLifetime or until a budget ends it.

maxLifetime is a hard deadline counted from creation, 24 hours by default. When it passes, the Sandbox is deleted whatever it is doing, running or stopped. It is the backstop that keeps a forgotten Sandbox from living forever. Create a new Sandbox, or one with a longer maxLifetime, to carry on.

A Sandbox has no outbound network access unless you open it. spec.network.egress.policy is one of:

  • Deny (the default): nothing leaves the Sandbox, so code in it cannot download or upload anything.
  • Open: the Sandbox can reach the public internet. It still cannot reach other Sandboxes, the machine it runs on or Nodus’s own services.

A Sandbox accepts no inbound connections; you reach it only through commands and file requests.

A Sandbox is billed per second, from the moment its machine slot is reserved until it is released, at the rate of its shape. The rate is the Nodus machine price for the CPU, memory and disk you asked for, and it already includes Nodus’s 12.5 % (how pricing works). Four things to know:

  • Starting and stopping are billed. The seconds spent starting (or restoring /workspace) and the seconds spent saving a stop are part of the reservation, shown as their own lines on the bill.
  • Stopped time is free. A stopped Sandbox holds no compute, so no compute cost accrues. The saved /workspace counts as storage, which is metered separately from compute.
  • The rate is frozen per start. The rate a start shows is the rate that start pays, even if prices change while it runs.
  • Shared capacity has limits. New shared capacity starts for you at most four times an hour, and none starts for 30 minutes after capacity started for you goes unused; until your organization’s first purchase it starts only for Sandboxes of up to 3 vCPU and 11 GiB. Otherwise a Sandbox runs on a machine of its own, at that machine’s rate.

nodus get sb/hello shows the running total in the COST column. Set spec.maxCostUSD to cap one Sandbox: it stops when it has cost that much, and you can raise the cap with another create by name. Without a cap, your credit balance and any Budget on the project are the limits.