Skip to content

Nodus for kubectl users

View Markdown

If you know kubectl, you already know most of nodus. Every Nodus resource (Jobs, Sandboxes, Volumes, Secrets, InferenceEndpoints, Budgets and the rest) is a Kubernetes-style object in the nodus.dev API group, with metadata, spec and status, and the CLI speaks the same verbs over all of them.

Kubernetes Nodus
Cluster The Nodus API (https://api.nodus-compute.ai)
Namespace (-n) Project (-p; -n is accepted)
User and credentials An API key per org, kept in the OS keychain
kubeconfig context One context per org, in ~/.nodus/config
kubectl get pods nodus get jobs, nodus get sandboxes, nodus get all
kubectl exec nodus exec into a Job, Sandbox or Workspace; each command is recorded as a Process

Your org comes from the API key, so there is no org in a manifest. Projects are created in the console or with nodus create -f, and default always exists.

Terminal window
nodus get jobs -o wide # tables are rendered by the server
nodus get jobs -l team=nlp --field-selector status.phase=Running
nodus get job/train -o jsonpath='{.status.phase}'
nodus get jobs -o custom-columns=NAME:.metadata.name,GPU:.spec.resources.gpu
nodus get jobs -w # watch
nodus describe job/train # includes a Placement section and events
nodus apply -f job.yaml # client-side three-way merge
nodus diff -f job.yaml # what apply would change (exit 1 on differences)
nodus apply -f jobs/ --prune -l app=nightly # delete what was applied before and is gone now
nodus edit job/train # $EDITOR, saved as a merge patch
nodus patch job/train --type merge --patch '{"spec":{"maxCostUSD":"50"}}'
nodus label job/train tier=gold
nodus wait job/train --for=jsonpath='{.status.phase}'=Succeeded --timeout 30m
nodus delete job/train --wait
nodus logs -f job/train
nodus exec -it sb/dev -- bash
nodus shell sandbox/dev # create the Sandbox if needed, then exec -it bash
nodus events -w --for job/train # events as they are recorded
nodus port-forward job/train 6006
nodus explain job.spec.resources
nodus api-resources
nodus auth can-i create jobs

Short names work as in kubectl: sb for sandboxes, vol for volumes, sec for secrets, sa for service accounts. nodus api-resources lists them all.

  • Server dry-run returns a cost estimate. nodus apply -f job.yaml --dry-run=server -o estimate prints the estimated cost, start time and the hold that would be placed, without creating anything.
  • State verbs instead of scaling. suspend, resume and cancel apply to Jobs, Pipelines and Sweeps; start and stop to Sandboxes, Workspaces, Functions, Agents and InferenceEndpoints. They set spec.state.
  • Requests are annotations. nodus request restart sb/dev (or nodus rollout restart sb/dev) asks a controller to act once, recorded on the object.
  • Typed generators. nodus create sandbox dev --cpu 2, nodus create secret hf --from-literal HF_TOKEN=…, nodus create volume weights --size 200Gi, nodus create budget research --limit 2000. Add --dry-run=client -o yaml to print the manifest instead of creating it. Kinds that show a secret once print it alone on standard output: nodus create apikey ci --scopes jobs:write, nodus create enrollmenttoken --pool lab (the host installer) and nodus create token sa/ci. nodus create sshkey --from-file ~/.ssh/id_ed25519.pub adds a public key, or --generate makes the pair.
  • Outputs are files you download. nodus cp job/train:outputs/model ./model copies a declared output and checks its SHA-256 digest.
  • Merge patches only. Strategic merge patch and server-side apply are not supported; apply does the three-way merge in the client and records the applied manifest in a last-applied annotation on the object.
  • Exit codes. Every command exits 0, 1 on an error or 2 on a usage error, like kubectl. nodus run passes your command’s exit code through instead.

~/.nodus/config is a real kubeconfig whose user entry runs nodus auth token as an exec credential plugin, so kubectl and any client-go tool (k9s included) can read and write Nodus resources:

Terminal window
export KUBECONFIG=~/.nodus/config
kubectl get jobs.nodus.dev
kubectl apply -f job.yaml
kubectl get jobs.nodus.dev -w
kubectl wait jobs.nodus.dev/train --for=jsonpath='{.status.phase}'=Succeeded

Use the full jobs.nodus.dev resource name with kubectl, because it also knows the built-in batch/v1 Jobs.

Like kubectl, nodus runs any executable called nodus-<name> on your PATH as nodus <name>, so you can add your own commands.