Logs and metrics
View MarkdownEverything your program writes to stdout and stderr is kept as its log. You can follow it while the program runs,
read it back after it finishes, and filter it by time, attempt, index or rank. Usage samples (GPU, CPU and memory)
are kept beside it, for nodus top and the console charts.
Follow a run
Section titled “Follow a run”nodus run streams the log until the Job finishes. To attach to a Job that is already running, or to come back
after you detached:
$ nodus logs -f job/finetune-llama-f (--follow) prints what is already stored, then new lines as they are written, without repeating or skipping
any. Press Ctrl+C to stop following; the Job keeps running.
The same command works for Sandboxes, Workspaces, Functions, Agents, AgentRuns, Images and individual Attempts:
nodus logs sb/dev, nodus logs attempt/NAME.
Choose which lines
Section titled “Choose which lines”| Flag | Shows |
|---|---|
--tail 100 |
The last 100 lines, then follows with -f |
--since 10m |
Lines from the last 10 minutes (any duration: 30s, 2h) |
--timestamps |
Each line prefixed with the time it was written |
--attempt 2 |
One attempt of the Job; the default is the latest one |
--index 3, --all-indexes |
One index of an Indexed Job, or all of them |
--rank 1, --all-ranks |
One rank of a multi-node Job (Beta), or all of them prefixed with [r<rank>] |
--process 12 |
The output of one Process in a Sandbox or Workspace, such as a nodus exec session |
--tail and --since combine: --since 1h --tail 50 prints at most the last 50 lines of the last hour. Without
--attempt, a Job that was recovered shows its newest attempt; earlier attempts stay readable by number, which
helps when a recovery followed a crash you want to look at.
From Python, job.logs() returns the same text, and the MCP logs tool reads up to 2,000 lines at a time.
Logs after the run
Section titled “Logs after the run”Logs are stored as they are written, not only on the machine that ran the program. When a Job finishes, is
suspended, or loses its capacity, nodus logs job/NAME still prints the whole log, and nodus logs --attempt 1
reads an attempt whose machine is long gone.
Logs are kept for 30 days. When capacity disappears without warning, the last few seconds written before the loss
may be missing; everything before them is kept. Secret values you pass to the Job (secrets:) are replaced with
[redacted] before the log is stored or shown.
Logs are your program’s own output, exactly as it wrote it. Nodus’ own messages about scheduling and recovery are
Events (nodus describe job/NAME, nodus events), never mixed into your log. nodus events -w prints the
events already recorded and then each new one as it happens; add --for job/NAME to follow one object, and
-o json to get the full Event objects.
Watch usage with nodus top
Section titled “Watch usage with nodus top”nodus top shows what running work is using right now:
$ nodus top jobsNAME GPU GPU MEM CPU MEMORY $/H SPENDfinetune-llama 94% 71.2Gi 3800m 41.0Gi $2.49 $6.12eval-sweep-3 61% 18.5Gi 1200m 12.3Gi $0.80 $0.35$ nodus top sandbox devGPU is the average utilization across the Job’s GPUs; GPU memory, CPU and memory are totals. Samples are taken
every 15 to 30 seconds, so a value can be up to half a minute old. nodus top sandboxes, workspaces and
functions work the same way.
Charts and queries
Section titled “Charts and queries”The console’s Metrics tab charts the same samples over time. For your own dashboards and scripts, the history of one object is at:
GET /metrics/v1/namespaces/PROJECT/jobs/NAME/history?metric=gpu&since=6h&step=1mmetric is gpu, gpuMemory, cpu or memory, the resource can be jobs, sandboxes, workspaces or
functions, and the answer has the shape of a Prometheus query_range result. Instead of since, pass start
and end as RFC 3339 times.
For anything else, the project’s samples answer PromQL in the Prometheus HTTP API at
/metrics/v1/namespaces/PROJECT/api/v1/query and query_range, so Grafana and other Prometheus clients can use
it as a data source with your API key. The series are:
| Series | Labels | Value |
|---|---|---|
nodus_container_gpu_utilization_percent |
kind, name, uid, attempt, gpu |
Utilization of one GPU, 0 to 100 |
nodus_container_gpu_memory_bytes |
kind, name, uid, attempt, gpu |
Memory in use on one GPU |
nodus_container_cpu_millis |
kind, name, uid, attempt |
CPU in use, in thousandths of a core |
nodus_container_memory_bytes |
kind, name, uid, attempt |
Memory in use |
Every query is limited to the project in the path: a query that names another org or project is refused. For
example, max_over_time(nodus_container_gpu_utilization_percent{kind="Job",name="finetune-llama"}[1h]) gives the
peak utilization of each GPU over the last hour.
Training metrics such as loss and accuracy are separate: report them with nodus.log.metrics(step=..., loss=...)
and they appear in the Job’s status under status.progress.metrics.