Skip to content

Logs and metrics

View Markdown

Everything your program writes to stdout and stderr is kept as its log. You can follow it while the program runs, read it back after it finishes, and filter it by time, attempt, index or rank. Usage samples (GPU, CPU and memory) are kept beside it, for nodus top and the console charts.

nodus run streams the log until the Job finishes. To attach to a Job that is already running, or to come back after you detached:

Terminal window
$ nodus logs -f job/finetune-llama

-f (--follow) prints what is already stored, then new lines as they are written, without repeating or skipping any. Press Ctrl+C to stop following; the Job keeps running.

The same command works for Sandboxes, Workspaces, Functions, Agents, AgentRuns, Images and individual Attempts: nodus logs sb/dev, nodus logs attempt/NAME.

Flag Shows
--tail 100 The last 100 lines, then follows with -f
--since 10m Lines from the last 10 minutes (any duration: 30s, 2h)
--timestamps Each line prefixed with the time it was written
--attempt 2 One attempt of the Job; the default is the latest one
--index 3, --all-indexes One index of an Indexed Job, or all of them
--rank 1, --all-ranks One rank of a multi-node Job (Beta), or all of them prefixed with [r<rank>]
--process 12 The output of one Process in a Sandbox or Workspace, such as a nodus exec session

--tail and --since combine: --since 1h --tail 50 prints at most the last 50 lines of the last hour. Without --attempt, a Job that was recovered shows its newest attempt; earlier attempts stay readable by number, which helps when a recovery followed a crash you want to look at.

From Python, job.logs() returns the same text, and the MCP logs tool reads up to 2,000 lines at a time.

Logs are stored as they are written, not only on the machine that ran the program. When a Job finishes, is suspended, or loses its capacity, nodus logs job/NAME still prints the whole log, and nodus logs --attempt 1 reads an attempt whose machine is long gone.

Logs are kept for 30 days. When capacity disappears without warning, the last few seconds written before the loss may be missing; everything before them is kept. Secret values you pass to the Job (secrets:) are replaced with [redacted] before the log is stored or shown.

Logs are your program’s own output, exactly as it wrote it. Nodus’ own messages about scheduling and recovery are Events (nodus describe job/NAME, nodus events), never mixed into your log. nodus events -w prints the events already recorded and then each new one as it happens; add --for job/NAME to follow one object, and -o json to get the full Event objects.

nodus top shows what running work is using right now:

Terminal window
$ nodus top jobs
NAME GPU GPU MEM CPU MEMORY $/H SPEND
finetune-llama 94% 71.2Gi 3800m 41.0Gi $2.49 $6.12
eval-sweep-3 61% 18.5Gi 1200m 12.3Gi $0.80 $0.35
$ nodus top sandbox dev

GPU is the average utilization across the Job’s GPUs; GPU memory, CPU and memory are totals. Samples are taken every 15 to 30 seconds, so a value can be up to half a minute old. nodus top sandboxes, workspaces and functions work the same way.

The console’s Metrics tab charts the same samples over time. For your own dashboards and scripts, the history of one object is at:

GET /metrics/v1/namespaces/PROJECT/jobs/NAME/history?metric=gpu&since=6h&step=1m

metric is gpu, gpuMemory, cpu or memory, the resource can be jobs, sandboxes, workspaces or functions, and the answer has the shape of a Prometheus query_range result. Instead of since, pass start and end as RFC 3339 times.

For anything else, the project’s samples answer PromQL in the Prometheus HTTP API at /metrics/v1/namespaces/PROJECT/api/v1/query and query_range, so Grafana and other Prometheus clients can use it as a data source with your API key. The series are:

Series Labels Value
nodus_container_gpu_utilization_percent kind, name, uid, attempt, gpu Utilization of one GPU, 0 to 100
nodus_container_gpu_memory_bytes kind, name, uid, attempt, gpu Memory in use on one GPU
nodus_container_cpu_millis kind, name, uid, attempt CPU in use, in thousandths of a core
nodus_container_memory_bytes kind, name, uid, attempt Memory in use

Every query is limited to the project in the path: a query that names another org or project is refused. For example, max_over_time(nodus_container_gpu_utilization_percent{kind="Job",name="finetune-llama"}[1h]) gives the peak utilization of each GPU over the last hour.

Training metrics such as loss and accuracy are separate: report them with nodus.log.metrics(step=..., loss=...) and they appear in the Job’s status under status.progress.metrics.