# Jobs

> Run a command to completion on a GPU or CPU, watch it, and download what it wrote.

Source: https://nodus-platform-site.pages.dev/docs/guides/jobs/
Build revision: 211ad9f836655b1c3a2668c4693e442471f28614

A Job runs your command until it finishes. You say what it needs (a GPU type, memory, a time limit, a spending cap) and Nodus runs it on capacity that fits, streams its logs, collects its outputs and charges only for the time it used.

## Sign in

Terminal window

```console
$ pip install nodus-compute
$ nodus login
```

`nodus login` opens the console in your browser and stores a key for this machine.

## Submit a Job

The fastest way is `nodus run`. It uploads the current directory, starts the command and follows it until it exits, passing the exit code through:

Terminal window

```console
$ nodus run --gpu L4 --image nodus/pytorch -- python hello.py
```

hello.py

```python
import json
import os

import torch

name = torch.cuda.get_device_name(0)
print(f"Hello from {name}")

# Anything written under /nodus/outputs is collected when the Job succeeds.
report = {"gpu": name, "cuda": torch.version.cuda, "index": os.environ["NODUS_INDEX"]}
with open(os.path.join(os.environ["NODUS_OUTPUTS_DIR"], "report.json"), "w") as f:
    json.dump(report, f)
```

Before anything runs, `nodus run` prints the estimate: the expected cost to completion and when it should start. Add `--max-cost 5` to stop the Job if it would spend more than $5, and `--timeout 2h` to limit its wall time.

To keep a Job in version control, write it as a manifest and apply it:

job.yaml

```yaml
apiVersion: nodus.dev/v1
kind: Job
metadata:
  name: hello
spec:
  image: nodus/pytorch
  command:
    - python
    - -c
    - |
      import json, os, torch
      name = torch.cuda.get_device_name(0)
      print(f"Hello from {name}")
      with open(os.path.join(os.environ["NODUS_OUTPUTS_DIR"], "report.json"), "w") as f:
          json.dump({"gpu": name, "index": os.environ["NODUS_INDEX"]}, f)
  resources:
    gpu: L4
  timeout: 10m
  maxCostUSD: "0.25"
  outputs:
    - name: report
      path: /nodus/outputs/report.json
```

Terminal window

```console
$ nodus apply -f job.yaml
job.nodus.dev/hello created
```

Anything under `/nodus/outputs` is collected when the Job succeeds. Declare the paths you want to download by name under `outputs`. `kubectl apply -f job.yaml` works too once your kubeconfig points at Nodus.

You do not have to name a GPU. Set `resources.gpu.minMemory` instead of `resources.gpu.type` and Nodus picks any accelerator with at least that much memory per GPU. If you also leave out `minMemory`, the annotations `nodus.dev/model` (for example `meta-llama/Llama-3.1-8B`) and `nodus.dev/dataset-bytes`, which you set under `metadata.annotations`, set it for you: 24Gi for 7B and 8B models, 32Gi for 13B and 14B, 48Gi for 34B and 40B, 80Gi for 70B and 72B, and 80Gi for a dataset over 1 TiB. `nodus apply -f job.yaml --dry-run=server -o yaml` shows the floor it picked.

## Watch it

In the console, open **Runs**, then select a run. **Logs** shows reported progress, stored outputs and the latest recovery checkpoint alongside the log stream. Distributed and indexed runs can be narrowed to a worker rank or index. The stream shows one worker at a time, starting with rank 0 and index 0. It follows that worker’s latest attempt; it is not a history of every retry. Connection failures show an error and **Retry logs**, while **Download loaded logs** saves the lines currently loaded in the browser. **Run activity** shows scheduling and lifecycle events separately from program output.

The command field preserves quoted arguments such as `python -c "print(1 + 1)"`. Shell pipelines and redirects need an explicit shell command, for example `bash -lc 'python prepare.py && python train.py'`.

Terminal window

```console
$ nodus get job/hello -w
$ nodus logs job/hello -f
$ nodus describe job/hello
```

`get` shows the phase, the GPU, the attempt, the cost so far and the age (`-o wide` adds the offering, its rate and progress). `describe` adds the conditions, the events and the reason for a failure with a suggested fix. A Job moves through these phases:

|Phase|What is happening|
|-|-|
|`Queued`|Waiting for capacity that fits and for a funded hold|
|`Provisioning`|Capacity is acquired; the image, source and inputs are being prepared|
|`Running`|Your command is running|
|`Recovering`|The capacity was lost; Nodus is moving the Job to new capacity|
|`Suspending`, `Suspended`|Saving state and releasing compute, then paused|
|`Cancelling`|Stopping and releasing compute|
|`Succeeded`, `Failed`, `Cancelled`|Finished; nothing is running or billed|

[Lifecycles](https://nodus-platform-site.pages.dev/docs/concepts/lifecycles/) lists every transition.

## Get the results

Open **Outputs** to download stored files and inspect their size and SHA-256. Distributed outputs identify their rank, and indexed outputs identify their index. A failed or cancelled run may have no stored outputs; its logs and activity explain what happened. **Details** includes attempts and the execution cost breakdown.

Terminal window

```console
$ nodus cp job/hello:outputs/report ./report.json
./report.json: 214 bytes, sha256 4f1c0e9a2b7d
```

`nodus cp` checks every download against the SHA-256 recorded when the Job finished. Outputs stay available until you delete the Job. [Outputs](https://nodus-platform-site.pages.dev/docs/guides/outputs/) covers directories, Indexed Jobs and loading results into a database.

## Run many indexes

An Indexed Job runs the same command `completions` times, at most `parallelism` at once. Each run sees its number in `NODUS_INDEX` (and in `JOB_COMPLETION_INDEX`, as on Kubernetes):

job.yaml

```yaml
apiVersion: nodus.dev/v1
kind: Job
metadata:
  name: indexed-outputs
spec:
  image: python:3.12-slim
  # Four indexes, two at a time. Each index sees its number in NODUS_INDEX.
  completions: 4
  parallelism: 2
  command:
    - python
    - -c
    - |
      import json, os
      i = int(os.environ["NODUS_INDEX"])
      rows = [{"index": i, "n": n, "square": n * n} for n in range(i * 100, (i + 1) * 100)]
      os.makedirs("/nodus/outputs/shards", exist_ok=True)
      with open(f"/nodus/outputs/shards/part-{i}.jsonl", "w") as f:
          f.writelines(json.dumps(r) + "\n" for r in rows)
      print(f"index {i}: wrote {len(rows)} rows")
  resources:
    cpu: "2"
    memory: 4Gi
  backoffLimit: 2
  timeout: 15m
  maxCostUSD: "0.10"
  outputs:
    - name: shards
      path: /nodus/outputs/shards
```

The Job succeeds when every index has succeeded. `backoffLimit` is how many failed runs the whole Job tolerates before it fails; a retried index starts fresh on new capacity. `status.completedIndexes` lists the finished indexes, such as `0-2,5`.

## Limit time and cost

|Field|Flag|What it does|
|-|-|-|
|`maxCostUSD`|`--max-cost`|The most the Job may spend. At the cap it saves its state and becomes `Suspended` with reason `MaxCostReached`; raise the cap to resume it. The cap can only be raised.|
|`timeout`|`--timeout`|Wall-clock limit counted from the first `Provisioning`; the Job fails with `DeadlineExceeded`. Time spent `Suspended` does not count. You can raise, lower or remove it on a running Job, and `status.timeoutTime` moves with it. `activeDeadlineSeconds` is accepted as an alias.|
|`completeByTime`|`--complete-by`|When you need the result. Nodus picks capacity that should finish in time.|
|`expectedDuration`|`--expected-duration`|Your estimate of the run time, used for the cost estimate before the Job has any history.|
|`placement.queueTimeout`||How long to wait for capacity before failing with `CapacityUnavailable`. The wait starts again when a suspended Job resumes.|

When credits run out or a budget is reached, a Job saves its state and becomes `Suspended` with reason `InsufficientCredits` or `BudgetExceeded`. It stops by `status.stopByTime`, the time its funds run out, even if the save is not finished, and resumes on its own once it is funded again.

## Suspend, resume and cancel

Terminal window

```console
$ nodus suspend job/train     # save state, release compute, stop billing
$ nodus resume job/train      # continue from the saved state
$ nodus cancel job/train      # stop for good
```

These set `spec.state` to `Suspended`, `Running` or `Cancelled`, so the same change works from a manifest. If a suspend cannot save state, the Job goes back to `Running` and its `Suspended` condition says `SuspendFailed`; nothing is lost. Nodus tries again ten minutes later, up to three tries in all (`status.suspendFailures` counts the failures), and then leaves the Job running. A Job with `recovery.continuity: Ephemeral` keeps no state, so suspending it restarts it from the beginning, and `nodus suspend` warns first. `nodus delete job/x` cancels a running Job before removing it.

## Recovery

Capacity can be lost while a Job runs, for example when interruptible capacity is reclaimed. What happens next depends on `recovery.continuity`:

|Continuity|After a loss|
|-|-|
|`Checkpointed` (default)|Resumes on new capacity from the last saved state in `NODUS_CHECKPOINT_DIR` (`/nodus/state`)|
|`Restartable`|Starts again from the beginning on new capacity|
|`Ephemeral`|Fails; nothing is kept|

Your program writes and loads its own state files in `NODUS_CHECKPOINT_DIR`; Nodus saves and restores that directory. Set `recovery.onInterruption: Fail` to fail instead of recovering. An index that uses up `recovery.maxAttempts` fails the Job with `RecoveryLimitExceeded`.

## When a Job fails

`status.reason` says why and `status.fix` says what to change. The common reasons:

|Reason|Meaning|
|-|-|
|`BackoffLimitExceeded`|Your command exited non-zero more times than `backoffLimit` allows|
|`OOMKilled`|The command ran out of memory; request more `resources.memory` or a larger GPU|
|`DeadlineExceeded`|`timeout` elapsed|
|`CapacityUnavailable`|No capacity fit the request within `placement.queueTimeout`|
|`ImagePullFailed`|The image does not exist or its registry refused the pull|
|`InvalidOutputPath`|A declared output was not written|
|`NoProgress`|Two recoveries in a row each ran less than `recovery.minProgressDuration` before losing capacity|
|`StorageUnavailable`|Nodus could not issue the storage credentials your command needs; run the Job again later|

## Clean up

Finished Jobs are deleted `ttlSecondsAfterFinished` seconds after they finish. Jobs from `nodus run` default to 30 days; pass `--keep` to keep one. Deleting a Job deletes its outputs.

## Multi-node Jobs

Beta

`spec.distributed` runs one Job across several nodes as a gang, with the rank, world size and rendezvous address in the environment. The Multi-node training guide covers launchers, networking and outputs from each rank.

## Environment

Every run sees these variables. Your own `env` cannot use the `NODUS_` prefix, except `NODUS_PARAM_`.

|Variable|Value|
|-|-|
|`NODUS_JOB`|The Job’s name|
|`NODUS_INDEX`, `JOB_COMPLETION_INDEX`|This run’s index, from 0|
|`NODUS_COMPLETIONS`|`spec.completions`|
|`NODUS_OUTPUTS_DIR`|`/nodus/outputs`|
|`NODUS_CHECKPOINT_DIR`|The first checkpoint path, `/nodus/state` by default|
|`NODUS_INPUT_<NAME>`|Where the input `<name>` is mounted, under `/nodus/inputs`|
|`NODUS_PARAM_<NAME>`|A Sweep parameter (see [Sweeps](https://nodus-platform-site.pages.dev/docs/guides/sweeps/))|

## Run code from GitHub

Install a GitHub Connection in the Job’s project and wait until it is Ready. Set `source.git.repo` to `owner/repository` or its GitHub HTTPS URL, and set `ref` to a branch, tag or commit. Nodus records the resolved commit and source archive in `status.source` before placing the Job. Retries and parallel indexes use the same files even when the branch moves.

```yaml
source:
  git:
    repo: your-org/your-repository
    ref: main
```

The files appear in `workingDir` before the command starts. Set `source.git.path` to use another checkout directory; this does not change the command’s working directory. The checkout contains repository files, without `.git` history, submodule contents or Git LFS downloads. GitHub credentials stay outside the guest. The source archive is limited to 100 MiB compressed, 512 MiB expanded and 10,000 entries. A missing Connection or rejected archive prevents execution.
