# Data and recovery

---

Images, volumes, outputs, secrets and checkpoint recovery.

---

# Checkpoints

> Keep a Job's progress through suspends, stops and lost capacity by saving it to the state directory, and answer checkpoint requests so every save is consistent.

Source: https://nodus-platform-site.pages.dev/docs/guides/checkpoints/
Build revision: 211ad9f836655b1c3a2668c4693e442471f28614

A checkpoint is a copy of your Job’s **state directory** that Nodus takes while the Job runs. When the Job is suspended, or the capacity under it is reclaimed, the next attempt starts with that directory restored and your program carries on from what it saved. Nodus decides when to checkpoint and where to store it; your program decides what to write and how to load it.

Checkpoints hold files, not memory. A restored attempt starts your command again from the beginning, with the state directory as it was at the last checkpoint, so the program must read its own progress back.

## Save your progress to the state directory

Write everything you need to continue, such as model weights, optimizer state and the current step, under `/nodus/state`. The path is also in `NODUS_STATE_DIR` (`NODUS_CHECKPOINT_DIR` is the older name for the same directory). On start, load what is there:

examples/checkpoints/resume/train.py

```python
"""A training loop that saves its progress where Nodus checkpoints it and resumes from it.

It uses only the standard library, so it runs in any image. With the Nodus SDK installed,
`nodus.state_dir()`, `nodus.checkpoint.on_request()` and `nodus.restored()` do the same work.
"""

import json
import os
import socket
import threading
import time
from pathlib import Path

STATE_DIR = Path(os.environ.get("NODUS_STATE_DIR", "/nodus/state"))
STATE = STATE_DIR / "progress.json"
STEPS = int(os.environ.get("STEPS", "120"))
SAVE_EVERY = 20  # like Hugging Face Trainer's save_steps: a regular save even when Nodus does not ask

pending = threading.Event()  # set while Nodus waits for a consistent checkpoint
request_seq = None
events = None


def save(step):
    """Write the state atomically, so a snapshot never sees a half-written file."""
    STATE_DIR.mkdir(parents=True, exist_ok=True)
    tmp = STATE.with_suffix(".tmp")
    tmp.write_text(json.dumps({"step": step}))
    os.replace(tmp, STATE)


def listen():
    """Subscribe to checkpoint requests on the events socket and flag each one for the loop."""
    global events, request_seq
    path = os.environ.get("NODUS_EVENTS_SOCKET", "/run/nodus/events.sock")
    try:
        events = socket.socket(socket.AF_UNIX, socket.SOCK_STREAM)
        events.connect(path)
    except OSError:
        return  # outside Nodus there is nobody to ask for checkpoints
    events.sendall(b'{"type":"checkpoint.subscribe"}\n')
    for line in events.makefile("r"):
        message = json.loads(line)
        if message.get("type") == "checkpoint.request":
            request_seq = message["seq"]
            pending.set()


def ack():
    """Tell Nodus the files in the state directory are complete."""
    events.sendall((json.dumps({"type": "checkpoint.ack", "seq": request_seq}) + "\n").encode())
    pending.clear()


def main():
    start = 0
    if STATE.exists():
        start = json.loads(STATE.read_text())["step"]
        print(f"resumed from step {start}", flush=True)
    threading.Thread(target=listen, daemon=True).start()
    for step in range(start + 1, STEPS + 1):
        time.sleep(1)  # one step of work
        print(f"step {step}/{STEPS}", flush=True)
        if pending.is_set():
            save(step)
            ack()
        elif step % SAVE_EVERY == 0:
            save(step)
    print("training complete", flush=True)


if __name__ == "__main__":
    main()
```

Two habits keep checkpoints useful:

* **Write atomically.** Write to a temporary file and rename it over the old one, as `save()` does. A checkpoint can then never capture a half-written file.
* **Keep outputs separate.** Results you want to download go to `/nodus/outputs`. The state directory is recovery state: it is restored into the next attempt, not offered as a download.

`NODUS_RESTORED=1` is set in an attempt that started from a checkpoint, if your program wants to log the difference. An empty state directory never replaces an earlier checkpoint that had files in it, so an attempt that fails before it saves anything cannot erase progress.

## Choose what is saved

`recovery.checkpoint` in the Job spec controls the checkpoint. The defaults suit most programs:

```yaml
recovery:
  continuity: Checkpointed      # restore the latest checkpoint into each new attempt
  checkpoint:
    paths: [/nodus/state]       # up to 64 absolute paths, all saved in one checkpoint
    interval: auto              # or a fixed interval from 1m to 6h
    maxSize: 1Ti                # larger checkpoints fail with CheckpointTooLarge
    retainAfterFinish: 168h     # after the Job finishes, keep only the final checkpoint
```

On the command line, `nodus run --checkpoint /nodus/state` sets `paths`. Saving a whole folder such as your working directory is possible by listing it, but it makes every checkpoint larger and slower; list only what you need to continue.

`continuity: Restartable` is the lighter choice for programs that track their position as a counter: Nodus keeps `NODUS_CURSOR_COMPLETED` and `NODUS_CURSOR_TOTAL` from your progress reports and skips the file copy. `continuity: Ephemeral` starts every attempt from scratch.

## When Nodus checkpoints

With `interval: auto`, Nodus sets the cadence from how often the capacity your Job runs on is interrupted and how long a checkpoint takes to save. A Job gets at least four checkpoints over its expected run time, and checkpointing takes no more than about a tenth of it. Capacity that is rarely interrupted is checkpointed less often.

Nodus also checkpoints, whatever the interval:

* when you run `nodus suspend job/NAME`, before the compute is released, so `nodus resume` continues from there;
* when the capacity gives notice that it is about to be reclaimed, so the replacement attempt loses as little work as possible.

`nodus describe job/NAME` shows the latest checkpoint, and `GET …/jobs/NAME/checkpoints` lists each one with its sequence number, attempt, time, size and file count.

## Answer checkpoint requests

A checkpoint taken while your program is halfway through writing its files would restore a broken state. To avoid that, a program can ask to be told before each checkpoint and say when its files are complete. This is the **request/ack handshake**, and the example above implements it with the standard library.

It runs over the events socket at `/run/nodus/events.sock` (`NODUS_EVENTS_SOCKET`), one JSON object per line:

1. Your program connects and sends `{"type": "checkpoint.subscribe"}` once.
2. Before each checkpoint, Nodus sends `{"type": "checkpoint.request", "seq": 7, "urgent": false}`.
3. Your program finishes the current step, writes its state and replies `{"type": "checkpoint.ack", "seq": 7}`.
4. Nodus copies the state directory, then your program carries on. It does not need to pause while the copy runs.

`urgent: true` means the capacity is about to go away: save at the next safe point and skip optional work. Nodus waits for the ack for as long as the shutdown allows and then takes the checkpoint anyway, so a program that hangs cannot block a suspend.

You choose whether to use the handshake with `recovery.checkpoint.integration`:

|Value|Behaviour|
|-|-|
|`Auto` (default)|Use the handshake when the program subscribes; otherwise checkpoint without asking|
|`None`|Never ask; checkpoint the paths as they are|
|`HFTrainer`|Answer requests from inside Hugging Face Trainer, with no change to your image ([below](https://nodus-platform-site.pages.dev/docs/guides/checkpoints/#hugging-face-trainer))|

With the Python SDK installed (`pip install nodus-compute`), `nodus.checkpoint.on_request(save)` registers a callback and sends the ack after it returns, `nodus.checkpoint.requested()` lets a loop poll instead, and `nodus.state_dir()` returns the directory. All of them do nothing outside Nodus, so the same script runs on your laptop.

## Hugging Face Trainer

Trainer already saves and resumes; point it at the state directory and it works with Nodus checkpoints:

* set `output_dir` to the state directory (`os.environ["NODUS_STATE_DIR"]`), or a folder inside it;
* set `save_steps` to how often Trainer saves on its own, and `save_total_limit` (for example `2`) so old checkpoints do not fill the directory;
* call `trainer.train(resume_from_checkpoint=True)` when `NODUS_RESTORED` is `1`, and `trainer.train()` on a first start, because Trainer refuses to resume from an empty directory.

To also save when Nodus asks, add `nodus.checkpoint.HFTrainerCallback()` to the Trainer’s callbacks, or set `integration: HFTrainer` and Nodus registers the same callback for you, even in an image without the SDK. Trainer then saves at the end of the current step and Nodus checkpoints once the save is written.

## Gang checkpoints (Beta)

Beta

Multi-node Jobs (`distributed`) are Beta. Their checkpoints use a different format.

A Job with `distributed` set defaults to `recovery.checkpoint.format: Dcp`. Every rank writes its shard with `torch.distributed.checkpoint` to `$NODUS_CHECKPOINT_URI`, and rank 0 commits it, which `nodus.checkpoint.dcp.save` and `.load` do for you. After a restart, `$NODUS_RESTORE_URI` names the latest committed checkpoint. A rank that lost its place in the gang cannot commit, so a restored gang always loads a checkpoint every rank finished.

## Run the example

The example suspends the Job mid-run and resumes it; the second attempt prints `resumed from step N` and finishes the remaining steps:

Terminal window

```console
$ cd examples/checkpoints/resume
$ nodus run --name checkpoints-resume --cpu 2 --checkpoint /nodus/state -d -- python train.py
$ nodus suspend job/checkpoints-resume
$ nodus resume job/checkpoints-resume
$ nodus logs -f job/checkpoints-resume
resumed from step 12
step 13/120
…
```

Checkpoints are deleted with their Job. After a Job finishes, Nodus keeps only its final checkpoint once `retainAfterFinish` (seven days by default) has passed.


---

# Connections

> Connect databases, S3 buckets, Weights & Biases and GitHub once, verified, and use them from Jobs, Sandboxes, imports and outputs.

Source: https://nodus-platform-site.pages.dev/docs/guides/connections/
Build revision: 211ad9f836655b1c3a2668c4693e442471f28614

A Connection is an external system that Nodus talks to for you: a Postgres database for query imports and outputs, an S3 bucket for inputs, a Weights & Biases project for live run tracking, or your GitHub repositories for private sources. Its credentials live in a [Secret](https://nodus-platform-site.pages.dev/docs/guides/secrets/); the Connection says what they are for, and Nodus checks that they work before anything uses them.

## Create a Connection

Put the credentials in a Secret, then point the Connection at it:

Terminal window

```console
nodus secret create analytics-db --from-literal DATABASE_URL=postgres://reader:...@db.example.com:5432/app
```

connection.yaml

```yaml
apiVersion: nodus.dev/v1
kind: Connection
metadata:
  name: ex-connections-postgres
spec:
  type: Postgres
  secret: ex-connections-postgres     # holds DATABASE_URL
  scope: Read                         # verified: the role can read, and is not asked to write
```

Terminal window

```console
nodus apply -f connection.yaml
nodus wait connection/ex-connections-postgres --for=condition=Verified
nodus describe connection/ex-connections-postgres
```

|`type`|Secret keys|Settings|Verified by|
|-|-|-|-|
|`Postgres`|`DATABASE_URL`|`scope`: `Read`, `Write` or `ReadWrite`|Connecting, a query, and for write scopes the right to create tables|
|`Neon`, `Supabase`|`DATABASE_URL`|`scope`; `neon.branch`|The same checks, on the provider’s own host|
|`S3`|`ROLE_ARN`, or `AWS_ACCESS_KEY_ID` and `AWS_SECRET_ACCESS_KEY`|`s3.bucket`, `s3.region`, `s3.prefix`, `s3.endpoint`|Reaching the bucket with short-lived credentials|
|`WandB`|`WANDB_API_KEY`|`wandb.entity`, `wandb.project`, `wandb.live`|The key’s access to the entity|

Prefer an IAM role (`ROLE_ARN`) for S3: Nodus assumes it for minutes at a time, so no long-lived key is stored. Every S3 Connection has its own external id in `status.externalId`, and Nodus sends exactly that id each time it assumes the role, so a role your trust policy grants to one Connection can’t be used through any other Connection or organization. Create the Connection, read the id, and require it in the role’s trust policy:

Terminal window

```console
nodus get connection/datasets -o jsonpath='{.status.externalId}'
```

```json
{
  "Effect": "Allow",
  "Principal": { "AWS": "<the Nodus principal shown in the console>" },
  "Action": "sts:AssumeRole",
  "Condition": { "StringEquals": { "sts:ExternalId": "<status.externalId>" } }
}
```

The Connection shows `Failed` until the trust policy names the id; the next check, within the hour, turns it `Ready`. You don’t need `EXTERNAL_ID` in the Secret. If you set it, it must equal `status.externalId`, and any other value is refused.

A Connection is `Ready` once verified, and `status.egressHosts` lists the hosts it may reach. Nodus checks it again every day, and every hour after a failure. When a check fails the Connection shows `Failed`, the `Verified` condition says why, and anything that needs it fails with `ConnectionNotReady` until you fix the Secret.

## Use it

```yaml
spec:
  connections: [tracking]               # injects WANDB_* and allows its hosts through the egress policy
  inputs:
    - name: shards
      bucket: {uri: s3://acme-data/shards/, connection: datasets}
```

A `WandB` Connection with `wandb.live: true` sets `WANDB_API_KEY`, `WANDB_ENTITY`, `WANDB_PROJECT`, `WANDB_RUN_GROUP`, `WANDB_NAME` and `WANDB_RUN_ID` for the Job, so `wandb.init()` needs no arguments. Each index of the Job logs to one run, and a run that Nodus restarts after lost capacity resumes it rather than starting a new one. A Job with one index shows the run’s page in `status.links`. A variable you set yourself keeps your value; setting `WANDB_RUN_ID`, `WANDB_ENTITY` or `WANDB_PROJECT` yourself means Nodus doesn’t link the run. A Job that names a Connection that is missing or not `Ready` runs without it and gets a `ConnectionNotReady` Event.

## Connect GitHub

A `GitHub` Connection installs the Nodus GitHub App on your account or organization. It needs no Secret:

Terminal window

```console
nodus create connection github --type GitHub
nodus get conn github -o jsonpath='{.status.installURL}'
```

Open the link, choose the repositories to share and install the App. GitHub returns you to Nodus, which checks that you can see the installation, and the Connection becomes `Ready` with `status.github` naming the account and repositories. `GET …/connections/github/repositories?limit=100` pages through all of them.

A Sandbox’s `init.git` can then clone a private repository: Nodus resolves the branch or tag to a commit when you create the Sandbox and clones that commit. The clone uses a short-lived token that can only read that one repository, and the token never appears in the Sandbox’s environment or files. If someone uninstalls the App, the Connection shows `Failed` with `InstallationRemoved` and a fresh `status.installURL` to install it again.

A Volume can import from a bucket (`source.s3`) or a query (`source.connectionQuery`); see [Volumes](https://nodus-platform-site.pages.dev/docs/guides/volumes/#import-from-elsewhere). A Job or Sandbox reaches only the hosts its Connections were verified against.


---

# Images

> Run on the catalog images, build your own from a few steps or a Dockerfile, and use private registries.

Source: https://nodus-platform-site.pages.dev/docs/guides/images/
Build revision: 211ad9f836655b1c3a2668c4693e442471f28614

Every Job, Sandbox, Function and Agent runs in a container image. You can use a catalog image, any public or private registry image, or an Image that Nodus builds for you. When you submit work, Nodus resolves the image’s tag to a digest and records it, so a retry or a recovery runs exactly the same bytes.

## Catalog images

|Image|Contents|Default for|
|-|-|-|
|`nodus/python:3.12` (also `3.10`, `3.11`, `3.13`)|Debian slim, Python, uv, git|Jobs without a GPU|
|`nodus/pytorch:2.8-cuda12.8`|CUDA 12.8 runtime, Python 3.12, PyTorch 2.8|Jobs with a GPU|
|`nodus/agent-tools`|Python 3.12, Node 22, git, ripgrep, the Nodus SDK|Sandboxes and Agents|

Terminal window

```console
nodus get images -n nodus
nodus run --gpu L4 --image nodus/pytorch -- python -c "import torch; print(torch.cuda.get_device_name())"
```

Catalog images run as uid 1000 and include `/bin/sh`.

## Build availability

Current hosted deployments support catalog images and prebuilt public or private registry images. Image builds from steps, Dockerfiles and Sandbox snapshots are unavailable. Those requests return `ImageBuildUnavailable` with HTTP 503 before accepting a build or charging for one. Publish your image to a registry, then use its reference directly or create an Image with only `base` and no steps. Environment packages that need a pip build must instead name a prebuilt `package.image`.

Existing unstarted build requests report `Failed` with an explanation instead of waiting indefinitely. The build examples below describe the contract for a deployment with a configured build runner and registry.

## Build an Image

An Image starts from a `base` and adds `steps`: `aptInstall`, `pipInstall`, `uvPipInstall`, `uvSync`, `run`, `copy`, `env` and `workdir`.

image.yaml

```yaml
apiVersion: nodus.dev/v1
kind: Image
metadata:
  name: ex-images-build
spec:
  base: nodus/python:3.12
  steps:
    - aptInstall: [jq]
    - pipInstall: {packages: ["requests==2.32.5"]}
    - env: {APP_MODE: example}
```

Terminal window

```console
nodus apply -f image.yaml
nodus wait image/ex-images-build --for=condition=Ready
```

The build log streams from `GET …/images/{name}/log?follow=true`; the Python SDK prints it while it builds.

Run on it with `imageRef`:

job.yaml

```yaml
apiVersion: nodus.dev/v1
kind: Job
metadata:
  name: ex-images-build
spec:
  imageRef: {name: ex-images-build}     # runs the digest the Image built, pinned at admission
  command: [python, -c, "import os, requests; print('requests', requests.__version__, os.environ['APP_MODE'])"]
```

In Python:

```python
image = (nodus.Image.from_registry("nodus/pytorch:2.8-cuda12.8")
         .apt_install("git").pip_install("transformers==4.57.6", "peft")
         .env({"HF_HUB_ENABLE_HF_TRANSFER": "1"}))
```

Builds run on Nodus and are billed per vCPU-second at the CPU rate. An Image whose spec matches one your org has already built, on the same base digest, reuses that digest without building (`status.build.cached: true`). An Image with only a `base` is ready at once: it is that image’s digest.

You can also build from a Dockerfile, inline or uploaded with its context:

```yaml
spec:
  dockerfile:
    inline: |
      FROM nodus/python:3.12
      RUN pip install polars==1.33.1
```

Build-time credentials go in `buildSecrets`: they are mounted for the build only and never stored in a layer.

## Private registries

Create a `Registry` Secret ([Secrets](https://nodus-platform-site.pages.dev/docs/guides/secrets/#private-registries)) and name it under `imagePullSecrets`, on the Job or Sandbox for `image`, or on the Image for a private `base`.

When work runs on a provider that pulls containers itself, Nodus copies the image by digest into your org’s space in the Nodus registry first, so your registry credential never reaches the provider. Such providers need `/bin/sh` in the image; an image without it gets the `ProviderContainerNeedsShell` warning and runs elsewhere.

## Pull a built Image yourself

A built Image lives in the Nodus registry under your org; `status.reference` holds its full reference. Log in with any username and an API key as the password, then pull it:

Terminal window

```bash
REF=$(nodus get image ex-images-build -o jsonpath='{.status.reference}')
echo "$NODUS_API_KEY" | docker login "${REF%%/*}" -u nodus --password-stdin
docker pull "$REF"
```

Viewers, and keys without the `images:write` scope, can pull but not push.

## Status and errors

`nodus get image ex-images-build` shows `PHASE`, `DIGEST` and `SIZE`. A failed build has a reason:

|Reason or error|Meaning|
|-|-|
|`BaseNotFound`|The base does not exist, or its registry refused the pull Secret|
|`StepFailed`|A step exited non-zero; the build log shows which|
|`BuildTimeout`|The build ran past its time limit|
|`ImageNotFound`, `ImagePullFailed`|A Job’s image could not be resolved when you submitted it|
|`ImageNotReady`|A Job or Sandbox names an Image that has not finished building|


---

# Outputs

> Collect files from a Job, download them with a verified checksum, and load results into Postgres.

Source: https://nodus-platform-site.pages.dev/docs/guides/outputs/
Build revision: 211ad9f836655b1c3a2668c4693e442471f28614

An output is a file or directory a Job writes for you to download. Outputs are collected when the Job succeeds, checksummed, and kept until you delete the Job.

## Write outputs

Write anything you want to keep under `/nodus/outputs` (also in `NODUS_OUTPUTS_DIR`). Every file there is collected when the Job succeeds. To give a file or directory a name of your own, declare it:

job.yaml

```yaml
apiVersion: nodus.dev/v1
kind: Job
metadata:
  name: hello
spec:
  image: nodus/pytorch
  command:
    - python
    - -c
    - |
      import json, os, torch
      name = torch.cuda.get_device_name(0)
      print(f"Hello from {name}")
      with open(os.path.join(os.environ["NODUS_OUTPUTS_DIR"], "report.json"), "w") as f:
          json.dump({"gpu": name, "index": os.environ["NODUS_INDEX"]}, f)
  resources:
    gpu: L4
  timeout: 10m
  maxCostUSD: "0.25"
  outputs:
    - name: report
      path: /nodus/outputs/report.json
```

A declared `path` is a file or a directory under `/nodus/outputs`. Every file is collected as its own output, named by its path below `/nodus/outputs`, and a declared name is another name for the file or directory it points at. A Job can declare up to 32 outputs; names use lowercase letters, digits, `.`, `_` and `-`, and the name `outputs` and the `nodus.` prefix are reserved.

Outputs are separate from checkpoints: files in `NODUS_CHECKPOINT_DIR` are for resuming the Job, not for download.

## Download outputs

Terminal window

```console
$ nodus cp job/hello:outputs/report ./report.json
./report.json: 214 bytes, sha256 4f1c0e9a2b7d
$ nodus get job/hello -o jsonpath='{.status.outputs}'
```

`nodus cp` writes the file only after its SHA-256 matches the digest recorded when the output was collected. Name a file by its declared name (`report`) or by its path below `/nodus/outputs` (`report.json`); the files of a directory output are listed one by one in `status.outputs`.

For an Indexed Job, every index writes its own copy of each output. Pick one with `--index`:

Terminal window

```console
$ nodus cp job/indexed-outputs:outputs/shards/part-2.jsonl ./part-2.jsonl --index 2
```

### With the API

`GET /apis/nodus.dev/v1/namespaces/<project>/jobs/<name>/outputs` lists the collected outputs with their index, size and `sha256`. `GET …/outputs/<output>` answers `302` with a short-lived download URL and the digest in the `X-Nodus-SHA256` header; add `?index=<n>` when several indexes produced the output. Verify the digest of what you download. Errors are `Status` objects with a `reason` and a `fix`: `OutputNotFound` (404) when the Job committed no such output, `OutputIndexRequired` (400) when you need to pass `?index=`, and `Invalid` (422) for a malformed `index`.

## Stage outputs

In a [Pipeline](https://nodus-platform-site.pages.dev/docs/guides/pipelines/), a stage reads an earlier stage’s output as an input, mounted read-only at `/nodus/inputs/<name>` and named by `NODUS_INPUT_<NAME>`. Any Job can do the same with an `output` input that names another Job:

```yaml
inputs:
  - name: data
    output: {job: prepare, name: shards}
```

## Load outputs into Postgres

An output with a `sink` is loaded into a table in your database after the Job succeeds. The database is reached through a Postgres, Neon or Supabase Connection:

job.yaml

```yaml
apiVersion: nodus.dev/v1
kind: Job
metadata:
  name: output-sink
spec:
  image: python:3.12-slim
  command:
    - python
    - -c
    - |
      import json
      with open("/nodus/outputs/squares.jsonl", "w") as f:
          for n in range(1000):
              f.write(json.dumps({"n": n, "square": n * n}) + "\n")
  timeout: 10m
  maxCostUSD: "0.05"
  outputs:
    - name: squares
      path: /nodus/outputs/squares.jsonl
      # After the Job succeeds, the rows are loaded into this table through the Connection named analytics.
      sink:
        connection: analytics
        table: squares
        mode: Replace
```

* The output must be a `.csv`, `.jsonl` or `.parquet` file.
* `mode: Append` (the default) adds rows; `Replace` replaces the table’s rows.
* Each file can be up to 5 GB and 50 million rows, and each record up to 8 MiB.
* The load runs as part of the Job and is billed to it as CPU time.

`status.outputs[].sink` shows each load’s phase (`Pending`, `Loading`, `Loaded` or `Failed`) and the rows loaded, and the `SinksLoaded` condition turns true when every load has finished. To retry the failed loads of a finished Job:

Terminal window

```console
$ nodus request reload-sinks job/output-sink
```

## How long outputs last

Outputs stay downloadable until the Job is deleted, and deleting the Job deletes them. Jobs started with `nodus run` are deleted 30 days after they finish unless you pass `--keep`; set `ttlSecondsAfterFinished` on a manifest to choose your own retention.


---

# Secrets

> Store API tokens, credentials and registry logins once, and pass them to Jobs, Sandboxes, Functions and Agents as env vars and files.

Source: https://nodus-platform-site.pages.dev/docs/guides/secrets/
Build revision: 211ad9f836655b1c3a2668c4693e442471f28614

A Secret holds named values, such as an API token or a database URL, encrypted with a key that belongs to your org. You create it once and name it in the Jobs and Sandboxes that need it. Nodus never returns a value through the API, the console or the CLI, and it redacts secret values from logs and outputs.

## Create a Secret

Terminal window

```console
nodus secret create hf-token --from-literal HF_TOKEN=hf_xxx
nodus secret create app-env --from-env-file .env
nodus secret create registry --type Registry \
  --from-literal server=ghcr.io --from-literal username=ada --from-literal password=ghp_xxx
```

Key names start with a letter or `_`, use letters, digits and `_`, and may not start with `NODUS_`. A Secret holds up to 64 keys of at most 4 KiB each. `nodus get secret hf-token` shows the key names and the current version, never the values.

## Use it in a Job

List the Secret under `secrets` to receive every key as an env var and as a read-only file `/run/secrets/<secret>/<key>`, or pick one key under another name with `valueFrom.secretKeyRef`:

job.yaml

```yaml
apiVersion: nodus.dev/v1
kind: Job
metadata:
  name: ex-secrets-env
spec:
  image: nodus/python:3.12
  secrets: [ex-secrets-env]          # every key as an env var and a file under /run/secrets/ex-secrets-env/
  env:
    - name: TOKEN                    # one key under another name
      valueFrom: {secretKeyRef: {name: ex-secrets-env, key: API_TOKEN}}
  command: [python, check.py]
```

check.py

```python
import os

token = os.environ["API_TOKEN"]
with open("/run/secrets/ex-secrets-env/API_TOKEN") as f:
    from_file = f.read()

# Print facts about the value, never the value: Nodus redacts secret values from logs anyway.
print(f"token has {len(token)} characters")
print("file matches" if from_file == token else "file differs")
print("alias matches" if os.environ["TOKEN"] == token else "alias differs")
```

Terminal window

```console
nodus apply -f job.yaml
nodus logs job/ex-secrets-env
```

The same `secrets` and `env` fields work on Sandboxes, Workspaces, Functions and Agents. In Python:

```python
hf = nodus.Secret.from_name("hf-token")
env = nodus.Secret.from_dotenv(".env")
```

A command whose literal `env` value equals one of the Secret values in its project is refused with `SecretValueInEnv`: reference the Secret instead of pasting its value. The same check applies to commands you start in a running Sandbox, Workspace or Job with `nodus exec` or the Processes API, which also refuse an `env` that sets the same name twice. Values shorter than 8 characters, and the `server` and `username` of a Registry Secret, are not checked.

A Job receives the Secret versions pinned when you submitted it, and a Sandbox the versions pinned when it started. If a Secret it uses is deleted (even if you create a new one with the same name), or a new value replaced a Job’s pinned version before a retry started, that attempt fails with `LaunchFailed`. Secret values delivered as env vars may total at most 1 MiB per container, and as files at most 8 MiB.

## Change a value

Writing a Secret again creates a new version; writing the same values again changes nothing. A Sandbox pins the newest version each time it starts, so it gets the new value on its next start, while a running Sandbox keeps the version it started with. A Job keeps the versions pinned when you submitted it, so every attempt runs with the same values; a retry that starts after you replace a value fails with `LaunchFailed`, so submit the Job again to use it. Earlier versions are deleted 7 days after they are replaced.

hf-token.yaml

```yaml
apiVersion: nodus.dev/v1
kind: Secret
metadata: {name: hf-token}
spec:
  stringData: {HF_TOKEN: hf_new}
```

Terminal window

```console
nodus apply -f hf-token.yaml
```

## Private registries

A Secret with `type: Registry` holds `server`, `username` and `password`. Name it under `imagePullSecrets` to run or build from a private image:

```yaml
spec:
  image: ghcr.io/acme/trainer:1.4
  imagePullSecrets: [{name: registry}]
```

Nodus resolves the tag to a digest when you submit, using the Secret, and the machine that runs the work pulls that digest with the same Secret version. The credential is used for the pull alone: it is not an env var or a file in the container. When the work runs on a provider that pulls containers itself, Nodus first copies the image by digest into your org’s space in the Nodus registry, so your registry credential never leaves Nodus.

## Limits and errors

|Error|Meaning|Fix|
|-|-|-|
|`SecretValueInEnv`|An `env` value equals a Secret value|Use `valueFrom.secretKeyRef` or `secrets`|
|`EncryptionFailed`|The value could not be encrypted|Retry; the Secret keeps its previous version|
|`ImagePullFailed`|The registry refused the pull Secret|Check `server`, `username` and `password`|


---

# Volumes

> Keep datasets, model weights and working files in Volumes, upload and download them, import from Hugging Face, git, URLs or S3, and mount them in Jobs and Sandboxes.

Source: https://nodus-platform-site.pages.dev/docs/guides/volumes/
Build revision: 211ad9f836655b1c3a2668c4693e442471f28614

A Volume is named storage that outlives the work that uses it. Every change is saved as a numbered revision, so you can see what changed, mount an earlier revision read-only, and never lose files to a stopped machine. You pay for the bytes stored, after your org’s included 10 GB.

## Create a Volume and upload files

volume.yaml

```yaml
apiVersion: nodus.dev/v1
kind: Volume
metadata:
  name: ex-volumes-put-get
spec:
  accessMode: ReadWriteOnce
  size: 1Gi
  source: {upload: {}}
```

Terminal window

```console
nodus apply -f volume.yaml                        # or: nodus create volume data --size 100Gi
nodus volume put ex-volumes-put-get ./data /data  # uploads ./data and prints the new revision
nodus volume ls ex-volumes-put-get /data
nodus volume get ex-volumes-put-get /data/hello.txt ./out/
nodus volume put data ./archive.tar.gz /raw --extract
nodus volume rm data /raw/old.csv
nodus volume clear data                           # an empty revision; the Volume and its history stay
```

Uploads only send what changed: a second `put` of the same files moves almost nothing. An upload that started from an older revision than the latest is refused with `409 Conflict`, so two people uploading at once never overwrite each other silently; run it again on top of the new revision.

`put` adds to what is there: a directory’s contents merge into the remote path, and a file lands at the remote path, or inside it when the remote path is a directory or ends in `/`. `rm` and `clear` also commit new revisions, so earlier revisions keep the removed files. `get` and `ls` read the latest revision, or the one `--revision <n>` names; `get … -` prints a file to standard output, and existing local files are kept unless `--force`. `--extract` expands a `.zip`, `.tar`, `.tar.gz`, `.tgz` or `.tar.zst` archive. It refuses an archive whose paths or links lead outside the remote path, or that has more than 10,000 members.

## Mount it in a Job

job.yaml

```yaml
apiVersion: nodus.dev/v1
kind: Job
metadata:
  name: ex-volumes-put-get
spec:
  image: nodus/python:3.12
  volumes:
    - {volume: ex-volumes-put-get, mountPath: /mnt/vol, readOnly: true}
  command: [cat, /mnt/vol/data/hello.txt]
```

|`accessMode`|Who can write|How changes are saved|
|-|-|-|
|`ReadWriteOnce` (default)|One attempt at a time|On exit and every `commitInterval` (default 5 minutes)|
|`ReadOnlyMany` (default for imports)|Nobody; any number of readers|Revisions come from imports and uploads|
|`ReadWriteMany`|Every worker|Each worker’s changes are published when it commits or exits; the last writer of a path wins|

A `ReadWriteOnce` Volume is held by one attempt at a time. Starting a second writer, or deleting the Volume while it is held, fails with `409 VolumeBusy`, which names the holder; mount it with `readOnly: true` to read alongside. Functions and Agents with more than one worker need `ReadOnlyMany` or `ReadWriteMany`.

Mount an earlier revision with `revision: <n>` and `readOnly: true`. `nodus get volume data` shows the latest `REVISION`, `USED` and the current `HOLDER`; `GET …/volumes/{name}/revisions` lists the kept revisions (`revisionHistoryLimit`, default 10).

## Import from elsewhere

Set exactly one `source`; the import runs once on Nodus, billed as CPU time, and produces revision 1:

```yaml
spec:
  accessMode: ReadOnlyMany
  size: 200Gi
  maxCostUSD: "2.00"
  source: {huggingface: {repo: meta-llama/Llama-3.1-8B, revision: 0e9e39f249a16976918f6564b8830bc894c89659, secret: hf-token}}
```

|Source|Fields|
|-|-|
|`huggingface`|`repo`, `revision` (a commit or tag), `files` globs, `secret` holding `HF_TOKEN`|
|`git`|`repo`, `ref`, `lfs`|
|`url`|an `https://` `url`, `sha256`, `extract: Auto` unpacks `.zip`, `.tar`, `.tar.gz`, `.tgz` and `.tar.zst`|
|`s3`|`uri` and an `S3` [Connection](https://nodus-platform-site.pages.dev/docs/guides/connections/)|

The Volume is `Pending` while it imports and `Ready` after. A failed import shows `Failed` with reason `ImportFailed` and a message; fix the source or add credits, then run `nodus request reimport volume/<name>`.

## In Python

```python
vol = nodus.Volume.from_name("data", create_if_missing=True)
vol.put_file("./local.csv", "/train/local.csv")
print(vol.listdir("/train"))
```
