Container runtime contract
View MarkdownEvery Job, Sandbox and Function worker runs your image in an isolated container. The same contract holds on every offering: your code can rely on the paths and variables below wherever Nodus places it.
Directories
Section titled “Directories”| Path | Variable | What it is for |
|---|---|---|
/nodus/state |
NODUS_STATE_DIR |
Recovery state. Write your model, optimizer and progress files here; Nodus checkpoints this directory and restores it on the next attempt. |
/nodus/inputs/<name> |
NODUS_INPUT_<NAME> |
Each declared input, read-only, downloaded before your command starts. The variable names the file (a URL or one object) or the directory (an object prefix, such as an earlier stage’s outputs). An input may set its own path instead. |
/nodus/outputs |
NODUS_OUTPUT_DIR |
Results. Every regular file here is uploaded when your command exits with code 0, with its SHA-256, and appears as the outputs output. Symbolic links are not collected. |
/dev/shm |
Shared memory for NCCL between GPUs and for data-loader workers: half the container’s memory limit, or half the machine’s memory without one. It counts against the memory limit. | |
/etc/nodus/hostfile |
Multi-node Jobs (Beta): one <address> slots=<GPUs per node> line per node in rank order, for DeepSpeed and MPI tools. |
|
/run/secrets/<name>/<key> |
Secret values as read-only files (mode 0400) on a memory-backed filesystem, so they never reach a disk. The environment also carries each value, named by its key. |
|
/run/nodus/events.sock |
NODUS_EVENTS_SOCKET |
Progress, metrics and the checkpoint handshake (below). |
/run/nodus/api.sock |
NODUS_RUNTIME_SOCKET |
The Nodus API and model calls, authenticated as your Job, Sandbox or Function (below). |
NODUS_CHECKPOINT_DIR is kept as another name for NODUS_STATE_DIR for existing programs.
Environment variables
Section titled “Environment variables”| Variable | Set for | Meaning |
|---|---|---|
NODUS_ATTEMPT |
Every container | The attempt id. A retried or recovered run gets a new attempt. |
NODUS_JOB, NODUS_INDEX, JOB_COMPLETION_INDEX |
Jobs | The Job and, for indexed Jobs, this index. |
NODUS_RESTORED |
Recovered attempts | 1 when /nodus/state was restored from a checkpoint before your command started. |
NODUS_CURSOR_COMPLETED, NODUS_CURSOR_TOTAL |
Restartable Jobs | The progress cursor your program last reported. |
NODUS_PARAM_<NAME> |
Sweep cells | The cell’s parameters. |
PET_NPROC_PER_NODE |
Single-node GPU Jobs | The GPU count, which torchrun reads as its --nproc-per-node default, so torchrun train.py starts one process per GPU. Your own value or flag wins. |
Restoring a checkpoint restores files, never process memory: your program starts from the beginning and reads
its own progress files from /nodus/state. Nodus never adds resume flags to your command.
Network
Section titled “Network”Jobs, Functions and Workspaces reach the public Internet by default; Sandboxes and Agents have no outbound
network unless their spec opens it. Either way a container never reaches private addresses (10.0.0.0/8,
172.16.0.0/12, 192.168.0.0/16), cloud metadata services, other containers or the machine it runs on, and has
no IPv6. localhost always works inside the container. /etc/resolv.conf points at public resolvers, and
/etc/hosts resolves localhost.
With an allow list (egress: AllowList) the container has no direct route out. HTTPS_PROXY and HTTP_PROXY
point at a proxy on the container’s loopback (http://127.0.0.1:3128) that reaches only the listed hosts, on
ports 80 and 443. A *.example.com entry allows every subdomain of example.com. Most HTTP clients (pip,
npm, curl, git over HTTPS, the Python and Node SDKs) use the proxy on their own. A listed name that
resolves to a private address is still refused.
CPU Sandboxes run under gVisor, whose localhost belongs to the sandbox alone, so there the proxy and the
inference proxy listen on the sandbox’s gateway address instead of 127.0.0.1. Read the address from
HTTPS_PROXY and OPENAI_BASE_URL rather than writing 127.0.0.1 into your code; NO_PROXY already covers it.
On a multi-node Job whose nodes share a private network, each rank’s eth0 lists the node’s private address
first, so NCCL, Gloo and torchrun advertise an address the other ranks reach. Ranks reach each other only on the
rendezvous and NCCL ports, whatever the Job’s outbound setting.
A container on a GPU offering sees exactly the GPUs assigned to it, numbered from 0 as CUDA and PyTorch see them,
with the NVIDIA driver libraries mounted read-only. It never sees another container’s GPUs. Use nvidia-smi or
torch.cuda.device_count() to check what you have; do not set CUDA_VISIBLE_DEVICES yourself.
Sidecars and the init command
Section titled “Sidecars and the init command”Sidecars start before your command, in the order you list them, in the same container: they share its files,
network, environment and logs. Each must answer its readiness probe (an HTTP GET of its path, any status below
400, or a TCP connection to its port) within 15 minutes before the next starts, and they stop when your command
exits. A sidecar runs from your container’s image.
initCommand runs after the sidecars and before your command, for preflight checks such as imports, free disk or
the GPU count. Your command starts only if it exits with code 0; any other code fails the attempt with that code,
and running past its timeout (30 minutes unless you set one) fails it with code 124. Both run inside your billed
time.
Commands you run in a container
Section titled “Commands you run in a container”Commands you start in a running container (nodus exec, Processes) run with the same environment, secrets and
directories as your main command, plus any variables you add. Their output is streamed to you and, for Processes,
kept with your logs with secret values masked. A command started from a terminal session stops when you
disconnect; a Process keeps running until it exits, you cancel it, its timeout passes or the container stops.
Port forwarding and preview URLs reach a server listening on the container’s localhost, whatever its outbound
network setting. In a CPU Sandbox they reach the sandbox’s own address instead, so listen on all interfaces
(0.0.0.0); the same holds for a sidecar’s readiness port. File operations (nodus cp, the console Files tab)
read, write and watch paths as your container sees them, including /nodus/state and /nodus/outputs, with the
permissions of your container’s user: files written this way belong to that user and appear only once complete.
/proc, /sys and /dev are not available to them.
The events socket
Section titled “The events socket”The socket speaks newline-delimited JSON, one object of at most 16 KiB per line. Send telemetry as objects with a
type and, to make retries safe, a unique id:
{"id": "evt-10", "type": "progress", "completed": 120, "total": 1000}{"id": "evt-11", "type": "log.metrics", "step": 1200, "loss": 0.41}When a program cannot reach the socket, it can print the same object on standard output after the prefix
nodus.event . The line also stays in your logs.
A program that sends no log.metrics events still gets its training metrics charted: Nodus reads loss, learning
rate, epoch, step and eval_* values from Hugging Face Trainer and PyTorch Lightning log lines until the program
sends a log.metrics event of its own. With recovery.checkpoint.integration: HFTrainer, the Trainer helper
package is mounted at /.nodus/python and put first on PYTHONPATH, so its sitecustomize registers the Nodus
callback in images without the Nodus SDK.
Checkpoint handshake
Section titled “Checkpoint handshake”Nodus decides when to checkpoint. To save a consistent state first, send
{"type": "checkpoint.subscribe"} once. Before each checkpoint Nodus writes a request on the connection:
{"type": "checkpoint.request", "requestId": "ck-1790000000000", "seq": 1790000000000, "urgent": true}Finish writing your files to /nodus/state, then answer with
{"type": "checkpoint.ready", "requestId": "ck-1790000000000"}. urgent means the capacity is about to go
away: save quickly. A program that never subscribes is checkpointed without being asked, so write your state
files atomically (write to a temporary name, then rename).
A checkpoint of an empty state directory never replaces an earlier checkpoint that had files in it.
The API socket
Section titled “The API socket”/run/nodus/api.sock is HTTP over a Unix socket. Requests under /apis/nodus.dev/ reach the Nodus API, and
requests under /v1/ (OpenAI- and Anthropic-compatible routes) reach Nodus inference. Both carry the service
account token of the Job, Sandbox or Function the container belongs to: Nodus adds it to each request, so the token is never in your environment or files, and
any Authorization header you send is replaced.
curl --unix-socket "$NODUS_RUNTIME_SOCKET" http://nodus/apis/nodus.dev/v1/...With the inference proxy enabled, the same model routes are also served on http://127.0.0.1:7777 inside the
container (on the gateway address in a CPU Sandbox), even with outbound network off, and OPENAI_BASE_URL, ANTHROPIC_BASE_URL, OPENAI_API_KEY and
ANTHROPIC_API_KEY are set for it. The keys are placeholders, so SDKs and tools that take a base URL work
unchanged. That port serves nothing but model calls.
Signals and exit codes
Section titled “Signals and exit codes”Your command runs under a small init process (tini, at /.nodus/bin) that forwards signals to your command’s
process group and reaps finished child processes, so your command does not need to be written as PID 1.
A stop sends SIGTERM to your command, then SIGKILL after the stop grace period (30 s unless your spec sets
another). Exit code 0 completes the attempt; any other code, or a signal, fails it with your exit code and the
last lines of your logs, with secret values masked.