Jobs
View MarkdownA Job runs your command until it finishes. You say what it needs (a GPU type, memory, a time limit, a spending cap) and Nodus runs it on capacity that fits, streams its logs, collects its outputs and charges only for the time it used.
Sign in
Section titled “Sign in”$ pip install nodus-compute$ nodus loginnodus login opens the console in your browser and stores a key for this machine.
Submit a Job
Section titled “Submit a Job”The fastest way is nodus run. It uploads the current directory, starts the command and follows it until it
exits, passing the exit code through:
$ nodus run --gpu L4 --image nodus/pytorch -- python hello.pyimport jsonimport os
import torch
name = torch.cuda.get_device_name(0)print(f"Hello from {name}")
# Anything written under /nodus/outputs is collected when the Job succeeds.report = {"gpu": name, "cuda": torch.version.cuda, "index": os.environ["NODUS_INDEX"]}with open(os.path.join(os.environ["NODUS_OUTPUTS_DIR"], "report.json"), "w") as f: json.dump(report, f)Before anything runs, nodus run prints the estimate: the expected cost to completion and when it should start.
Add --max-cost 5 to stop the Job if it would spend more than $5, and --timeout 2h to limit its wall time.
To keep a Job in version control, write it as a manifest and apply it:
apiVersion: nodus.dev/v1kind: Jobmetadata: name: hellospec: image: nodus/pytorch command: - python - -c - | import json, os, torch name = torch.cuda.get_device_name(0) print(f"Hello from {name}") with open(os.path.join(os.environ["NODUS_OUTPUTS_DIR"], "report.json"), "w") as f: json.dump({"gpu": name, "index": os.environ["NODUS_INDEX"]}, f) resources: gpu: L4 timeout: 10m maxCostUSD: "0.25" outputs: - name: report path: /nodus/outputs/report.json$ nodus apply -f job.yamljob.nodus.dev/hello createdAnything under /nodus/outputs is collected when the Job succeeds. Declare the paths you want to download by name
under outputs. kubectl apply -f job.yaml works too once your kubeconfig points at Nodus.
You do not have to name a GPU. Set resources.gpu.minMemory instead of resources.gpu.type and Nodus picks any
accelerator with at least that much memory per GPU. If you also leave out minMemory, the annotations
nodus.dev/model (for example meta-llama/Llama-3.1-8B) and nodus.dev/dataset-bytes, which you set under
metadata.annotations, set it for you: 24Gi for 7B and 8B models, 32Gi for 13B and 14B, 48Gi for 34B and 40B,
80Gi for 70B and 72B, and 80Gi for a dataset over 1 TiB.
nodus apply -f job.yaml --dry-run=server -o yaml shows the floor it picked.
Watch it
Section titled “Watch it”In the console, open Runs, then select a run. Logs shows reported progress, stored outputs and the latest recovery checkpoint alongside the log stream. Distributed and indexed runs can be narrowed to a worker rank or index. The stream shows one worker at a time, starting with rank 0 and index 0. It follows that worker’s latest attempt; it is not a history of every retry. Connection failures show an error and Retry logs, while Download loaded logs saves the lines currently loaded in the browser. Run activity shows scheduling and lifecycle events separately from program output.
The command field preserves quoted arguments such as python -c "print(1 + 1)". Shell pipelines and redirects
need an explicit shell command, for example bash -lc 'python prepare.py && python train.py'.
$ nodus get job/hello -w$ nodus logs job/hello -f$ nodus describe job/helloget shows the phase, the GPU, the attempt, the cost so far and the age (-o wide adds the offering, its rate and
progress). describe adds the conditions, the events and the reason for a failure with a suggested fix. A Job moves through these phases:
| Phase | What is happening |
|---|---|
Queued |
Waiting for capacity that fits and for a funded hold |
Provisioning |
Capacity is acquired; the image, source and inputs are being prepared |
Running |
Your command is running |
Recovering |
The capacity was lost; Nodus is moving the Job to new capacity |
Suspending, Suspended |
Saving state and releasing compute, then paused |
Cancelling |
Stopping and releasing compute |
Succeeded, Failed, Cancelled |
Finished; nothing is running or billed |
Lifecycles lists every transition.
Get the results
Section titled “Get the results”Open Outputs to download stored files and inspect their size and SHA-256. Distributed outputs identify their rank, and indexed outputs identify their index. A failed or cancelled run may have no stored outputs; its logs and activity explain what happened. Details includes attempts and the execution cost breakdown.
$ nodus cp job/hello:outputs/report ./report.json./report.json: 214 bytes, sha256 4f1c0e9a2b7dnodus cp checks every download against the SHA-256 recorded when the Job finished. Outputs stay available until
you delete the Job. Outputs covers directories, Indexed Jobs and loading results into a
database.
Run many indexes
Section titled “Run many indexes”An Indexed Job runs the same command completions times, at most parallelism at once. Each run sees its number
in NODUS_INDEX (and in JOB_COMPLETION_INDEX, as on Kubernetes):
apiVersion: nodus.dev/v1kind: Jobmetadata: name: indexed-outputsspec: image: python:3.12-slim # Four indexes, two at a time. Each index sees its number in NODUS_INDEX. completions: 4 parallelism: 2 command: - python - -c - | import json, os i = int(os.environ["NODUS_INDEX"]) rows = [{"index": i, "n": n, "square": n * n} for n in range(i * 100, (i + 1) * 100)] os.makedirs("/nodus/outputs/shards", exist_ok=True) with open(f"/nodus/outputs/shards/part-{i}.jsonl", "w") as f: f.writelines(json.dumps(r) + "\n" for r in rows) print(f"index {i}: wrote {len(rows)} rows") resources: cpu: "2" memory: 4Gi backoffLimit: 2 timeout: 15m maxCostUSD: "0.10" outputs: - name: shards path: /nodus/outputs/shardsThe Job succeeds when every index has succeeded. backoffLimit is how many failed runs the whole Job tolerates
before it fails; a retried index starts fresh on new capacity. status.completedIndexes lists the finished
indexes, such as 0-2,5.
Limit time and cost
Section titled “Limit time and cost”| Field | Flag | What it does |
|---|---|---|
maxCostUSD |
--max-cost |
The most the Job may spend. At the cap it saves its state and becomes Suspended with reason MaxCostReached; raise the cap to resume it. The cap can only be raised. |
timeout |
--timeout |
Wall-clock limit counted from the first Provisioning; the Job fails with DeadlineExceeded. Time spent Suspended does not count. You can raise, lower or remove it on a running Job, and status.timeoutTime moves with it. activeDeadlineSeconds is accepted as an alias. |
completeByTime |
--complete-by |
When you need the result. Nodus picks capacity that should finish in time. |
expectedDuration |
--expected-duration |
Your estimate of the run time, used for the cost estimate before the Job has any history. |
placement.queueTimeout |
How long to wait for capacity before failing with CapacityUnavailable. The wait starts again when a suspended Job resumes. |
When credits run out or a budget is reached, a Job saves its state and becomes Suspended with reason
InsufficientCredits or BudgetExceeded. It stops by status.stopByTime, the time its funds run out, even if the
save is not finished, and resumes on its own once it is funded again.
Suspend, resume and cancel
Section titled “Suspend, resume and cancel”$ nodus suspend job/train # save state, release compute, stop billing$ nodus resume job/train # continue from the saved state$ nodus cancel job/train # stop for goodThese set spec.state to Suspended, Running or Cancelled, so the same change works from a manifest. If a
suspend cannot save state, the Job goes back to Running and its Suspended condition says SuspendFailed;
nothing is lost. Nodus tries again ten minutes later, up to three tries in all (status.suspendFailures counts the
failures), and then leaves the Job running. A Job with recovery.continuity: Ephemeral keeps no state, so suspending it restarts it from the
beginning, and nodus suspend warns first. nodus delete job/x cancels a running Job before removing it.
Recovery
Section titled “Recovery”Capacity can be lost while a Job runs, for example when interruptible capacity is reclaimed. What happens next
depends on recovery.continuity:
| Continuity | After a loss |
|---|---|
Checkpointed (default) |
Resumes on new capacity from the last saved state in NODUS_CHECKPOINT_DIR (/nodus/state) |
Restartable |
Starts again from the beginning on new capacity |
Ephemeral |
Fails; nothing is kept |
Your program writes and loads its own state files in NODUS_CHECKPOINT_DIR; Nodus saves and restores that
directory. Set recovery.onInterruption: Fail to fail instead of recovering. An index that uses up
recovery.maxAttempts fails the Job with RecoveryLimitExceeded.
When a Job fails
Section titled “When a Job fails”status.reason says why and status.fix says what to change. The common reasons:
| Reason | Meaning |
|---|---|
BackoffLimitExceeded |
Your command exited non-zero more times than backoffLimit allows |
OOMKilled |
The command ran out of memory; request more resources.memory or a larger GPU |
DeadlineExceeded |
timeout elapsed |
CapacityUnavailable |
No capacity fit the request within placement.queueTimeout |
ImagePullFailed |
The image does not exist or its registry refused the pull |
InvalidOutputPath |
A declared output was not written |
NoProgress |
Two recoveries in a row each ran less than recovery.minProgressDuration before losing capacity |
StorageUnavailable |
Nodus could not issue the storage credentials your command needs; run the Job again later |
Clean up
Section titled “Clean up”Finished Jobs are deleted ttlSecondsAfterFinished seconds after they finish. Jobs from nodus run default to 30
days; pass --keep to keep one. Deleting a Job deletes its outputs.
Multi-node Jobs
Section titled “Multi-node Jobs”Environment
Section titled “Environment”Every run sees these variables. Your own env cannot use the NODUS_ prefix, except NODUS_PARAM_.
| Variable | Value |
|---|---|
NODUS_JOB |
The Job’s name |
NODUS_INDEX, JOB_COMPLETION_INDEX |
This run’s index, from 0 |
NODUS_COMPLETIONS |
spec.completions |
NODUS_OUTPUTS_DIR |
/nodus/outputs |
NODUS_CHECKPOINT_DIR |
The first checkpoint path, /nodus/state by default |
NODUS_INPUT_<NAME> |
Where the input <name> is mounted, under /nodus/inputs |
NODUS_PARAM_<NAME> |
A Sweep parameter (see Sweeps) |
Run code from GitHub
Section titled “Run code from GitHub”Install a GitHub Connection in the Job’s project and wait until it is Ready. Set
source.git.repo to owner/repository or its GitHub HTTPS URL, and set ref to a
branch, tag or commit. Nodus records the resolved commit and source archive in
status.source before placing the Job. Retries and parallel indexes use the same
files even when the branch moves.
source: git: repo: your-org/your-repository ref: mainThe files appear in workingDir before the command starts. Set source.git.path
to use another checkout directory; this does not change the command’s working
directory. The checkout contains repository files, without .git history,
submodule contents or Git LFS downloads. GitHub credentials stay outside the guest.
The source archive is limited to 100 MiB compressed, 512 MiB expanded and 10,000
entries. A missing Connection or rejected archive prevents execution.