Checkpoints, stored logs and nodus top
Jobs now keep their progress, and their logs, wherever they run.
- Checkpoints. Write your state to
/nodus/stateand a suspended or recovered Job continues from its latest checkpoint. Nodus sets the cadence from the capacity’s interruption rate, and an empty directory never replaces a checkpoint that had files in it. - The checkpoint handshake. Subscribe on the events socket and Nodus asks before each checkpoint, so your
program saves a consistent state first. Hugging Face Trainer works with
integration: HFTrainer, even in images without the SDK. - Logs after the run.
nodus logs job/NAMEreads the whole log after the Job finishes or its capacity is gone, for 30 days.-ffollows without repeating or skipping lines;--tail,--since,--attempt,--indexand--processpick the lines you want. nodus top. GPU, GPU memory, CPU and memory for running Jobs, Sandboxes, Workspaces and Functions, with the hourly rate and spend. History is in the console’s Metrics tab and at a Prometheus-compatible query API.
See Checkpoints and Logs and metrics.