# Sandbox isolation

> What keeps a Sandbox apart from the machine, from other Sandboxes and from the network, and what is kept when it stops.

Source: https://nodus-platform-site.pages.dev/docs/concepts/sandboxes-isolation/
Build revision: 211ad9f836655b1c3a2668c4693e442471f28614

A [Sandbox](https://nodus-platform-site.pages.dev/docs/guides/sandboxes/) runs code you did not write, such as the output of a model, so the boundary around it matters more than it does for your own jobs. This page lists the layers of that boundary, in the order a piece of code meets them.

## Where Sandboxes run

Sandboxes run on **CPU machines Nodus operates**, not on GPU machines rented for your jobs. Nodus packs machines **per organization**: a machine serves one organization at a time, your Sandboxes share it with your other Sandboxes, agent workers and CPU jobs, and with nobody else’s. An empty machine is destroyed after 10 minutes and never handed to another organization, so one organization’s files and memory never sit on a machine another organization uses next.

## Layers

### 1. A user-space kernel (gVisor)

Each Sandbox runs under [gVisor](https://gvisor.dev/) (`runsc`, in its `systrap` mode). gVisor answers the Sandbox’s system calls in its own kernel written in a memory-safe language and passes the host kernel only a small, filtered set of calls. A bug in the host kernel’s handling of an unusual system call is reachable from a Sandbox only through that narrow filter, not directly.

### 2. A network namespace and a firewall per Sandbox

Every Sandbox gets its own network namespace with a default-deny `nftables` firewall:

* Private ranges (RFC 1918), carrier-grade NAT, link-local addresses and cloud metadata addresses are always blocked, whatever the policy, so a Sandbox cannot reach the machine, its neighbours or the cloud’s metadata service.
* Policy `Deny` (the default) has no route out at all.
* Policy `Open` translates traffic to public addresses only.

Nothing is allowed in from the network. Bytes in and out are counted, and a blocked connection raises an `EgressDenied` Event on the Sandbox.

### 3. Resource limits

Each Sandbox runs in its own cgroup (v2) that limits CPU and memory to what you asked for and the number of processes to 256, so a fork bomb or a memory leak stops at the Sandbox’s own limits instead of slowing its neighbours. The root filesystem is backed by a per-Sandbox file under a disk quota (`resources.disk`), not by a directory shared with other Sandboxes.

### 4. An unprivileged process

Commands run as a non-root user with no Linux capabilities and `no_new_privs` set, so the process cannot gain privileges through a set-uid program. The user comes from the image. The Nodus helper that serves commands and files runs outside the user’s process tree and is read-only to it.

## Continuity: what survives a stop

`spec.continuity.mode` decides what a Sandbox keeps when it stops or when its machine is lost:

|Mode|A stop keeps|If the machine is lost|
|-|-|-|
|`Snapshotted` (default)|The filesystem changes the Sandbox made, `/workspace` first among them|The Sandbox restarts on another machine from its last snapshot; work since that snapshot is gone|
|`Ephemeral`|Nothing: the next start is empty|The Sandbox moves to `Failed` with the reason `NodeLost`|

Processes never survive a stop: a start runs a fresh container from the image with the saved files restored, so start long-lived servers again from your code. Keep your work in `/workspace`, the directory the Sandbox starts in. `Snapshotted` is the right choice for an agent that works in a repository; `Ephemeral` suits a fresh sandbox per task. Snapshots are taken at every stop and about every 10 minutes while the Sandbox runs.

## Where gVisor is not available

Production Nodus CPU nodes ship with `runsc`, and Sandboxes there run under gVisor as described above. On a node without `runsc` (the local development provider on a laptop, or a Linux host where gVisor is not installed), `nodusd` runs the Sandbox under `runc` instead, with the same network namespace, firewall, cgroup limits and unprivileged process. Layers 2 to 4 are identical; layer 1 is not: a `runc` Sandbox shares the host’s Linux kernel, so it is a boundary for development and trusted code, not for hostile code.

Caution

Do not run untrusted code on a development node that lacks `runsc`.

## What a Sandbox does not protect against

* **Secrets you put in it.** Anything in a Sandbox’s environment or files is readable by the code running there. With `egress.policy: Open`, that code can send it anywhere on the internet. Keep `Deny` unless the work needs the network, and give a Sandbox only the credentials its task needs.
* **Resource use up to your limits.** A Sandbox can use all the CPU and memory it asked for, and is billed for them while it runs. Set `maxCostUSD` to cap the spend of code you do not control.
