Skip to content

A Job runs your command until it finishes. You say what it needs (a GPU type, memory, a time limit, a spending cap) and Nodus runs it on capacity that fits, streams its logs, collects its outputs and charges only for the time it used.

Terminal window
$ pip install nodus-compute
$ nodus login

nodus login opens the console in your browser and stores a key for this machine.

The fastest way is nodus run. It uploads the current directory, starts the command and follows it until it exits, passing the exit code through:

Terminal window
$ nodus run --gpu L4 --image nodus/pytorch -- python hello.py
hello.py
import json
import os
import torch
name = torch.cuda.get_device_name(0)
print(f"Hello from {name}")
# Anything written under /nodus/outputs is collected when the Job succeeds.
report = {"gpu": name, "cuda": torch.version.cuda, "index": os.environ["NODUS_INDEX"]}
with open(os.path.join(os.environ["NODUS_OUTPUTS_DIR"], "report.json"), "w") as f:
json.dump(report, f)

Before anything runs, nodus run prints the estimate: the expected cost to completion and when it should start. Add --max-cost 5 to stop the Job if it would spend more than $5, and --timeout 2h to limit its wall time.

To keep a Job in version control, write it as a manifest and apply it:

job.yaml
apiVersion: nodus.dev/v1
kind: Job
metadata:
name: hello
spec:
image: nodus/pytorch
command:
- python
- -c
- |
import json, os, torch
name = torch.cuda.get_device_name(0)
print(f"Hello from {name}")
with open(os.path.join(os.environ["NODUS_OUTPUTS_DIR"], "report.json"), "w") as f:
json.dump({"gpu": name, "index": os.environ["NODUS_INDEX"]}, f)
resources:
gpu: L4
timeout: 10m
maxCostUSD: "0.25"
outputs:
- name: report
path: /nodus/outputs/report.json
Terminal window
$ nodus apply -f job.yaml
job.nodus.dev/hello created

Anything under /nodus/outputs is collected when the Job succeeds. Declare the paths you want to download by name under outputs. kubectl apply -f job.yaml works too once your kubeconfig points at Nodus.

You do not have to name a GPU. Set resources.gpu.minMemory instead of resources.gpu.type and Nodus picks any accelerator with at least that much memory per GPU. If you also leave out minMemory, the annotations nodus.dev/model (for example meta-llama/Llama-3.1-8B) and nodus.dev/dataset-bytes, which you set under metadata.annotations, set it for you: 24Gi for 7B and 8B models, 32Gi for 13B and 14B, 48Gi for 34B and 40B, 80Gi for 70B and 72B, and 80Gi for a dataset over 1 TiB. nodus apply -f job.yaml --dry-run=server -o yaml shows the floor it picked.

In the console, open Runs, then select a run. Logs shows reported progress, stored outputs and the latest recovery checkpoint alongside the log stream. Distributed and indexed runs can be narrowed to a worker rank or index. The stream shows one worker at a time, starting with rank 0 and index 0. It follows that worker’s latest attempt; it is not a history of every retry. Connection failures show an error and Retry logs, while Download loaded logs saves the lines currently loaded in the browser. Run activity shows scheduling and lifecycle events separately from program output.

The command field preserves quoted arguments such as python -c "print(1 + 1)". Shell pipelines and redirects need an explicit shell command, for example bash -lc 'python prepare.py && python train.py'.

Terminal window
$ nodus get job/hello -w
$ nodus logs job/hello -f
$ nodus describe job/hello

get shows the phase, the GPU, the attempt, the cost so far and the age (-o wide adds the offering, its rate and progress). describe adds the conditions, the events and the reason for a failure with a suggested fix. A Job moves through these phases:

Phase What is happening
Queued Waiting for capacity that fits and for a funded hold
Provisioning Capacity is acquired; the image, source and inputs are being prepared
Running Your command is running
Recovering The capacity was lost; Nodus is moving the Job to new capacity
Suspending, Suspended Saving state and releasing compute, then paused
Cancelling Stopping and releasing compute
Succeeded, Failed, Cancelled Finished; nothing is running or billed

Lifecycles lists every transition.

Open Outputs to download stored files and inspect their size and SHA-256. Distributed outputs identify their rank, and indexed outputs identify their index. A failed or cancelled run may have no stored outputs; its logs and activity explain what happened. Details includes attempts and the execution cost breakdown.

Terminal window
$ nodus cp job/hello:outputs/report ./report.json
./report.json: 214 bytes, sha256 4f1c0e9a2b7d

nodus cp checks every download against the SHA-256 recorded when the Job finished. Outputs stay available until you delete the Job. Outputs covers directories, Indexed Jobs and loading results into a database.

An Indexed Job runs the same command completions times, at most parallelism at once. Each run sees its number in NODUS_INDEX (and in JOB_COMPLETION_INDEX, as on Kubernetes):

job.yaml
apiVersion: nodus.dev/v1
kind: Job
metadata:
name: indexed-outputs
spec:
image: python:3.12-slim
# Four indexes, two at a time. Each index sees its number in NODUS_INDEX.
completions: 4
parallelism: 2
command:
- python
- -c
- |
import json, os
i = int(os.environ["NODUS_INDEX"])
rows = [{"index": i, "n": n, "square": n * n} for n in range(i * 100, (i + 1) * 100)]
os.makedirs("/nodus/outputs/shards", exist_ok=True)
with open(f"/nodus/outputs/shards/part-{i}.jsonl", "w") as f:
f.writelines(json.dumps(r) + "\n" for r in rows)
print(f"index {i}: wrote {len(rows)} rows")
resources:
cpu: "2"
memory: 4Gi
backoffLimit: 2
timeout: 15m
maxCostUSD: "0.10"
outputs:
- name: shards
path: /nodus/outputs/shards

The Job succeeds when every index has succeeded. backoffLimit is how many failed runs the whole Job tolerates before it fails; a retried index starts fresh on new capacity. status.completedIndexes lists the finished indexes, such as 0-2,5.

Field Flag What it does
maxCostUSD --max-cost The most the Job may spend. At the cap it saves its state and becomes Suspended with reason MaxCostReached; raise the cap to resume it. The cap can only be raised.
timeout --timeout Wall-clock limit counted from the first Provisioning; the Job fails with DeadlineExceeded. Time spent Suspended does not count. You can raise, lower or remove it on a running Job, and status.timeoutTime moves with it. activeDeadlineSeconds is accepted as an alias.
completeByTime --complete-by When you need the result. Nodus picks capacity that should finish in time.
expectedDuration --expected-duration Your estimate of the run time, used for the cost estimate before the Job has any history.
placement.queueTimeout How long to wait for capacity before failing with CapacityUnavailable. The wait starts again when a suspended Job resumes.

When credits run out or a budget is reached, a Job saves its state and becomes Suspended with reason InsufficientCredits or BudgetExceeded. It stops by status.stopByTime, the time its funds run out, even if the save is not finished, and resumes on its own once it is funded again.

Terminal window
$ nodus suspend job/train # save state, release compute, stop billing
$ nodus resume job/train # continue from the saved state
$ nodus cancel job/train # stop for good

These set spec.state to Suspended, Running or Cancelled, so the same change works from a manifest. If a suspend cannot save state, the Job goes back to Running and its Suspended condition says SuspendFailed; nothing is lost. Nodus tries again ten minutes later, up to three tries in all (status.suspendFailures counts the failures), and then leaves the Job running. A Job with recovery.continuity: Ephemeral keeps no state, so suspending it restarts it from the beginning, and nodus suspend warns first. nodus delete job/x cancels a running Job before removing it.

Capacity can be lost while a Job runs, for example when interruptible capacity is reclaimed. What happens next depends on recovery.continuity:

Continuity After a loss
Checkpointed (default) Resumes on new capacity from the last saved state in NODUS_CHECKPOINT_DIR (/nodus/state)
Restartable Starts again from the beginning on new capacity
Ephemeral Fails; nothing is kept

Your program writes and loads its own state files in NODUS_CHECKPOINT_DIR; Nodus saves and restores that directory. Set recovery.onInterruption: Fail to fail instead of recovering. An index that uses up recovery.maxAttempts fails the Job with RecoveryLimitExceeded.

status.reason says why and status.fix says what to change. The common reasons:

Reason Meaning
BackoffLimitExceeded Your command exited non-zero more times than backoffLimit allows
OOMKilled The command ran out of memory; request more resources.memory or a larger GPU
DeadlineExceeded timeout elapsed
CapacityUnavailable No capacity fit the request within placement.queueTimeout
ImagePullFailed The image does not exist or its registry refused the pull
InvalidOutputPath A declared output was not written
NoProgress Two recoveries in a row each ran less than recovery.minProgressDuration before losing capacity
StorageUnavailable Nodus could not issue the storage credentials your command needs; run the Job again later

Finished Jobs are deleted ttlSecondsAfterFinished seconds after they finish. Jobs from nodus run default to 30 days; pass --keep to keep one. Deleting a Job deletes its outputs.

Every run sees these variables. Your own env cannot use the NODUS_ prefix, except NODUS_PARAM_.

Variable Value
NODUS_JOB The Job’s name
NODUS_INDEX, JOB_COMPLETION_INDEX This run’s index, from 0
NODUS_COMPLETIONS spec.completions
NODUS_OUTPUTS_DIR /nodus/outputs
NODUS_CHECKPOINT_DIR The first checkpoint path, /nodus/state by default
NODUS_INPUT_<NAME> Where the input <name> is mounted, under /nodus/inputs
NODUS_PARAM_<NAME> A Sweep parameter (see Sweeps)

Install a GitHub Connection in the Job’s project and wait until it is Ready. Set source.git.repo to owner/repository or its GitHub HTTPS URL, and set ref to a branch, tag or commit. Nodus records the resolved commit and source archive in status.source before placing the Job. Retries and parallel indexes use the same files even when the branch moves.

source:
git:
repo: your-org/your-repository
ref: main

The files appear in workingDir before the command starts. Set source.git.path to use another checkout directory; this does not change the command’s working directory. The checkout contains repository files, without .git history, submodule contents or Git LFS downloads. GitHub credentials stay outside the guest. The source archive is limited to 100 MiB compressed, 512 MiB expanded and 10,000 entries. A missing Connection or rejected archive prevents execution.