# Logs and metrics

> Follow a Job's output live, read it back after the capacity is gone, and watch GPU, CPU and memory use with nodus top.

Source: https://nodus-platform-site.pages.dev/docs/guides/logs/
Build revision: 211ad9f836655b1c3a2668c4693e442471f28614

Everything your program writes to stdout and stderr is kept as its log. You can follow it while the program runs, read it back after it finishes, and filter it by time, attempt, index or rank. Usage samples (GPU, CPU and memory) are kept beside it, for `nodus top` and the console charts.

## Follow a run

`nodus run` streams the log until the Job finishes. To attach to a Job that is already running, or to come back after you detached:

Terminal window

```console
$ nodus logs -f job/finetune-llama
```

`-f` (`--follow`) prints what is already stored, then new lines as they are written, without repeating or skipping any. Press Ctrl+C to stop following; the Job keeps running.

The same command works for Sandboxes, Workspaces, Functions, Agents, AgentRuns, Images and individual Attempts: `nodus logs sb/dev`, `nodus logs attempt/NAME`.

## Choose which lines

|Flag|Shows|
|-|-|
|`--tail 100`|The last 100 lines, then follows with `-f`|
|`--since 10m`|Lines from the last 10 minutes (any duration: `30s`, `2h`)|
|`--timestamps`|Each line prefixed with the time it was written|
|`--attempt 2`|One attempt of the Job; the default is the latest one|
|`--index 3`, `--all-indexes`|One index of an Indexed Job, or all of them|
|`--rank 1`, `--all-ranks`|One rank of a multi-node Job (Beta), or all of them prefixed with `[r<rank>]`|
|`--process 12`|The output of one Process in a Sandbox or Workspace, such as a `nodus exec` session|

`--tail` and `--since` combine: `--since 1h --tail 50` prints at most the last 50 lines of the last hour. Without `--attempt`, a Job that was recovered shows its newest attempt; earlier attempts stay readable by number, which helps when a recovery followed a crash you want to look at.

From Python, `job.logs()` returns the same text, and the MCP `logs` tool reads up to 2,000 lines at a time.

## Logs after the run

Logs are stored as they are written, not only on the machine that ran the program. When a Job finishes, is suspended, or loses its capacity, `nodus logs job/NAME` still prints the whole log, and `nodus logs --attempt 1` reads an attempt whose machine is long gone.

Logs are kept for 30 days. When capacity disappears without warning, the last few seconds written before the loss may be missing; everything before them is kept. Secret values you pass to the Job (`secrets:`) are replaced with `[redacted]` before the log is stored or shown.

Logs are your program’s own output, exactly as it wrote it. Nodus’ own messages about scheduling and recovery are **Events** (`nodus describe job/NAME`, `nodus events`), never mixed into your log. `nodus events -w` prints the events already recorded and then each new one as it happens; add `--for job/NAME` to follow one object, and `-o json` to get the full Event objects.

## Watch usage with nodus top

`nodus top` shows what running work is using right now:

Terminal window

```console
$ nodus top jobs
NAME             GPU   GPU MEM   CPU     MEMORY   $/H     SPEND
finetune-llama   94%   71.2Gi    3800m   41.0Gi   $2.49   $6.12
eval-sweep-3     61%   18.5Gi    1200m   12.3Gi   $0.80   $0.35
$ nodus top sandbox dev
```

GPU is the average utilization across the Job’s GPUs; GPU memory, CPU and memory are totals. Samples are taken every 15 to 30 seconds, so a value can be up to half a minute old. `nodus top sandboxes`, `workspaces` and `functions` work the same way.

## Charts and queries

The console’s **Metrics** tab charts the same samples over time. For your own dashboards and scripts, the history of one object is at:

```text
GET /metrics/v1/namespaces/PROJECT/jobs/NAME/history?metric=gpu&since=6h&step=1m
```

`metric` is `gpu`, `gpuMemory`, `cpu` or `memory`, the resource can be `jobs`, `sandboxes`, `workspaces` or `functions`, and the answer has the shape of a Prometheus `query_range` result. Instead of `since`, pass `start` and `end` as RFC 3339 times.

For anything else, the project’s samples answer PromQL in the Prometheus HTTP API at `/metrics/v1/namespaces/PROJECT/api/v1/query` and `query_range`, so Grafana and other Prometheus clients can use it as a data source with your API key. The series are:

|Series|Labels|Value|
|-|-|-|
|`nodus_container_gpu_utilization_percent`|`kind`, `name`, `uid`, `attempt`, `gpu`|Utilization of one GPU, 0 to 100|
|`nodus_container_gpu_memory_bytes`|`kind`, `name`, `uid`, `attempt`, `gpu`|Memory in use on one GPU|
|`nodus_container_cpu_millis`|`kind`, `name`, `uid`, `attempt`|CPU in use, in thousandths of a core|
|`nodus_container_memory_bytes`|`kind`, `name`, `uid`, `attempt`|Memory in use|

Every query is limited to the project in the path: a query that names another org or project is refused. For example, `max_over_time(nodus_container_gpu_utilization_percent{kind="Job",name="finetune-llama"}[1h])` gives the peak utilization of each GPU over the last hour.

Training metrics such as loss and accuracy are separate: report them with `nodus.log.metrics(step=..., loss=...)` and they appear in the Job’s status under `status.progress.metrics`.
