Skip to content

The resource model

View Markdown

Everything you run on Nodus is an object: a Job, a Sandbox, a Function, a Volume, a Budget. Every object has the same shape and the same verbs, so once you know one kind you know them all, and the CLI, the Python SDK, the console, MCP clients and kubectl all work the same way.

apiVersion: nodus.dev/v1
kind: Job
metadata:
name: train-a # unique per project and kind
namespace: default # the project
labels:
app: trainer
spec: # what you want
image: nodus/pytorch:2.8-cuda12.8
command: [python, train.py]
maxCostUSD: "40.00"
status: # what Nodus observed; written by Nodus only
phase: Running
  • metadata.name is a DNS label of at most 63 characters, unique within its project and kind. Set metadata.generateName instead to get a unique suffix; such a create needs an Idempotency-Key header.
  • metadata.namespace is the project. Every org starts with the project default. The read-only project nodus holds what Nodus publishes, such as base images, which you reference as nodus/<name>:<tag>.
  • metadata.uid is a stable id such as job_01j9…, and metadata.resourceVersion changes on every write. Send the resourceVersion you read with an update to make it conditional: if someone changed the object in between, the update fails with error code Conflict and you read it again.
  • Labels select objects (-l app=trainer). Labels and annotations under the nodus.dev/ prefix are set by Nodus; the label nodus.dev/created-by records who created an object and nodus.dev/launched-by from where.

You change what an object does by changing its spec. Stopping, suspending and cancelling are values of spec.state, not separate actions: nodus suspend job/train-a sets spec.state: Suspended, and a later nodus apply of the same file keeps it suspended. Actions that change nothing in spec, such as restarting a Sandbox’s service, are requests: nodus request restart sandbox/dev sets the annotation nodus.dev/restart-requested-at to the current time, and Nodus acts once per value it has not handled yet.

Most spec fields cannot change after create. An update that changes one fails with error code FieldImmutable, and the response lists each field. The fields that can change, such as spec.state and a raise of spec.maxCostUSD, are listed in each kind’s reference.

status.phase uses the same words on every kind, so --field-selector status.phase=Running means the same thing everywhere. Run-to-completion kinds move through Queued, Provisioning, Running and end in Succeeded, Failed or Cancelled; long-running kinds move through Pending, Starting, Running and Stopped. status.conditions explain the details, such as the condition Ready.

  • Create by name is safe to retry. Creating an object that already exists with the same spec returns the existing object. A different spec fails with error code AlreadyExists and lists the differences.
  • Review before you pay. Add ?dryRun=All (--dry-run=server in the CLI) to run every check without creating anything. The response shows the object with its defaults and, for kinds that use compute, status.estimate: the expected cost, the first hold and the start time. Send the ETag of that response as If-Match on the real create to launch exactly what you reviewed: if the spec or the price book changed, the create fails with error code PreconditionFailed and you review again. If only the hourly rate moved, the create goes ahead and never runs above the reviewed rate plus 10 % (unless you set placement.maxRateUSDPerHour); when nothing fits, the object waits in Queued and its status says the price is above the estimate.
  • Retries never double-create. The CLI and SDK send an Idempotency-Key with every create. Retrying with the same key returns the first response with the header Idempotent-Replayed: true; reusing a key for a different request fails with error code IdempotencyKeyReused.

Every list can be watched: nodus get jobs -w, or ?watch=true&resourceVersion=<from a list> on the API. A watch delivers every change after that point, in order, and never skips one. Watches send newline-separated events (ADDED, MODIFIED, DELETED) as JSON, or as application/x-ndjson when you ask for it; with allowWatchBookmarks=true you also get a bookmark every 30 seconds to resume from. Nodus keeps 24 hours of changes: resuming from an older point fails with error code Expired, and you list again.

Lists return at most 500 objects by default (limit, up to 1,000) and a continue token for the next page.

nodus delete job/train-a stops the work first and removes the object once cleanup is done: while it waits the object shows a deletionTimestamp, and the phase Cancelling or Terminating. Nodus tracks that cleanup with finalizers such as the finalizer nodus.dev/billing, which clears when the final charge has posted, so you never see an object gone while it is still costing money. Deleting an object also deletes what it created (a Pipeline’s Jobs, for example); with --cascade=orphan those stay.

Every error has the same body: a Kubernetes Status whose reason is a stable error code, plus fix (what to do next), docs (a page for that code) and requestId (quote it to support). The error reference lists every code.

The API speaks the Kubernetes wire protocol for these objects, so kubectl, k9s and client-go work against it: projects are namespaces, kubectl get jobs.nodus.dev lists Jobs, and kubectl explain, apply, diff, wait and -w behave as they do on a cluster. The nodus CLI adds what kubectl does not have, such as logs, exec and file transfer for Nodus kinds.