# Outputs

> Collect files from a Job, download them with a verified checksum, and load results into Postgres.

Source: https://nodus-platform-site.pages.dev/docs/guides/outputs/
Build revision: 211ad9f836655b1c3a2668c4693e442471f28614

An output is a file or directory a Job writes for you to download. Outputs are collected when the Job succeeds, checksummed, and kept until you delete the Job.

## Write outputs

Write anything you want to keep under `/nodus/outputs` (also in `NODUS_OUTPUTS_DIR`). Every file there is collected when the Job succeeds. To give a file or directory a name of your own, declare it:

job.yaml

```yaml
apiVersion: nodus.dev/v1
kind: Job
metadata:
  name: hello
spec:
  image: nodus/pytorch
  command:
    - python
    - -c
    - |
      import json, os, torch
      name = torch.cuda.get_device_name(0)
      print(f"Hello from {name}")
      with open(os.path.join(os.environ["NODUS_OUTPUTS_DIR"], "report.json"), "w") as f:
          json.dump({"gpu": name, "index": os.environ["NODUS_INDEX"]}, f)
  resources:
    gpu: L4
  timeout: 10m
  maxCostUSD: "0.25"
  outputs:
    - name: report
      path: /nodus/outputs/report.json
```

A declared `path` is a file or a directory under `/nodus/outputs`. Every file is collected as its own output, named by its path below `/nodus/outputs`, and a declared name is another name for the file or directory it points at. A Job can declare up to 32 outputs; names use lowercase letters, digits, `.`, `_` and `-`, and the name `outputs` and the `nodus.` prefix are reserved.

Outputs are separate from checkpoints: files in `NODUS_CHECKPOINT_DIR` are for resuming the Job, not for download.

## Download outputs

Terminal window

```console
$ nodus cp job/hello:outputs/report ./report.json
./report.json: 214 bytes, sha256 4f1c0e9a2b7d
$ nodus get job/hello -o jsonpath='{.status.outputs}'
```

`nodus cp` writes the file only after its SHA-256 matches the digest recorded when the output was collected. Name a file by its declared name (`report`) or by its path below `/nodus/outputs` (`report.json`); the files of a directory output are listed one by one in `status.outputs`.

For an Indexed Job, every index writes its own copy of each output. Pick one with `--index`:

Terminal window

```console
$ nodus cp job/indexed-outputs:outputs/shards/part-2.jsonl ./part-2.jsonl --index 2
```

### With the API

`GET /apis/nodus.dev/v1/namespaces/<project>/jobs/<name>/outputs` lists the collected outputs with their index, size and `sha256`. `GET …/outputs/<output>` answers `302` with a short-lived download URL and the digest in the `X-Nodus-SHA256` header; add `?index=<n>` when several indexes produced the output. Verify the digest of what you download. Errors are `Status` objects with a `reason` and a `fix`: `OutputNotFound` (404) when the Job committed no such output, `OutputIndexRequired` (400) when you need to pass `?index=`, and `Invalid` (422) for a malformed `index`.

## Stage outputs

In a [Pipeline](https://nodus-platform-site.pages.dev/docs/guides/pipelines/), a stage reads an earlier stage’s output as an input, mounted read-only at `/nodus/inputs/<name>` and named by `NODUS_INPUT_<NAME>`. Any Job can do the same with an `output` input that names another Job:

```yaml
inputs:
  - name: data
    output: {job: prepare, name: shards}
```

## Load outputs into Postgres

An output with a `sink` is loaded into a table in your database after the Job succeeds. The database is reached through a Postgres, Neon or Supabase Connection:

job.yaml

```yaml
apiVersion: nodus.dev/v1
kind: Job
metadata:
  name: output-sink
spec:
  image: python:3.12-slim
  command:
    - python
    - -c
    - |
      import json
      with open("/nodus/outputs/squares.jsonl", "w") as f:
          for n in range(1000):
              f.write(json.dumps({"n": n, "square": n * n}) + "\n")
  timeout: 10m
  maxCostUSD: "0.05"
  outputs:
    - name: squares
      path: /nodus/outputs/squares.jsonl
      # After the Job succeeds, the rows are loaded into this table through the Connection named analytics.
      sink:
        connection: analytics
        table: squares
        mode: Replace
```

* The output must be a `.csv`, `.jsonl` or `.parquet` file.
* `mode: Append` (the default) adds rows; `Replace` replaces the table’s rows.
* Each file can be up to 5 GB and 50 million rows, and each record up to 8 MiB.
* The load runs as part of the Job and is billed to it as CPU time.

`status.outputs[].sink` shows each load’s phase (`Pending`, `Loading`, `Loaded` or `Failed`) and the rows loaded, and the `SinksLoaded` condition turns true when every load has finished. To retry the failed loads of a finished Job:

Terminal window

```console
$ nodus request reload-sinks job/output-sink
```

## How long outputs last

Outputs stay downloadable until the Job is deleted, and deleting the Job deletes them. Jobs started with `nodus run` are deleted 30 days after they finish unless you pass `--keep`; set `ttlSecondsAfterFinished` on a manifest to choose your own retention.
