Attempts and recovery
View MarkdownEvery run on Nodus executes as one or more attempts. An attempt is one incarnation of your command on one machine. When that machine is reclaimed, loses its network or fails, Nodus starts a new attempt somewhere else and your run continues from what it saved. This page explains what carries over, how Nodus decides a machine is gone, and what you pay for along the way.
Attempts and epochs
Section titled “Attempts and epochs”A Job index, a Sandbox or a worker slot holds one running attempt at a time. Each new attempt gets the next epoch, a number that only grows. Nodus accepts reports, checkpoints and outputs only from the current epoch, so a machine that comes back after it was replaced can never overwrite the work of its successor.
You see a run’s attempts with nodus get attempts -l nodus.dev/job=<name>. A run’s attempts share its
name with an epoch suffix, and gang members add a rank suffix (-r1, -r2).
| Attempt phase | Meaning |
|---|---|
Pending, Placing, Acquiring |
Choosing and preparing capacity |
Starting |
The machine is ready; the image is pulled and inputs or a checkpoint are restored |
Running |
Your command is running |
Succeeded |
Your command exited 0 |
Failed |
Your command or its machine failed; reason says which (NodeLost, Preempted, OOMKilled, …) |
Cancelled |
The attempt was stopped on purpose: a suspend, a cancel, a budget or lifetime limit |
Continuity modes
Section titled “Continuity modes”recovery.continuity says what a new attempt starts from:
| Mode | A new attempt starts with | Use it for |
|---|---|---|
Checkpointed |
The latest committed checkpoint of your declared state paths (NODUS_CHECKPOINT_DIR by default) |
Training and long jobs that save their own model, optimizer and progress files |
Restartable |
A cold start plus the progress cursor you reported (NODUS_CURSOR_COMPLETED, NODUS_CURSOR_TOTAL) |
Batch work that can skip what it already finished |
Ephemeral |
A cold start | Short or idempotent work |
Snapshotted |
The latest filesystem snapshot | Sandboxes |
Restoring files restores files only, never process memory. Your program loads its own checkpoint when it starts; Nodus never adds resume flags to your command.
An empty checkpoint never counts as saved progress and never replaces an earlier useful one. If every checkpoint a
run commits is empty, the run shows Checkpointed=False, reason=NotCheckpointable, and Nodus plans and prices it
as Ephemeral from then on.
What survives a preemption
Section titled “What survives a preemption”When a provider reclaims interruptible capacity or a machine stops answering, Nodus recovers make-before-break:
- On a reclaim notice, Nodus asks your attempt for an urgent checkpoint and, at the same time, starts preparing a replacement machine.
- The replacement is prepared up to the point where it could start, but it does not start while the old machine might still be writing.
- It starts only once the old machine is provably gone: the old attempt acknowledged its stop, the provider confirmed the machine terminated, or the old attempt’s lease ran out.
- The new attempt restores according to the continuity mode above.
So a Checkpointed run loses at most the work since its last committed checkpoint, and a checkpoint committed
during the reclaim notice still counts. If the old machine comes back before it is replaced, the replacement is
released and your run simply continues; you are not charged for the replacement.
A machine that stops sending heartbeats gets a replacement prepared after 20 seconds. After 60 seconds of silence an attempt on shared capacity is declared lost and replaced. An attempt on a machine dedicated to your organization may keep running up to the edge of its funding while its replacement waits, because such a machine can only lose its own work.
Recovery limits
Section titled “Recovery limits”Recovery stops, and the run fails with a typed reason, when:
| Reason | Rule |
|---|---|
NoProgress |
Two attempts in a row made no progress. Progress means the command started and either ran for recovery.minProgressDuration (default 2 minutes) or committed a non-empty checkpoint. A preemption after progress never counts |
RestoreFailed |
The same checkpoint failed to restore twice |
ImagePullFailed |
The image failed to pull on a second machine (a pull failure is retried once elsewhere) |
MaxAttemptsExceeded |
recovery.maxAttempts recoveries were used (default 8; 3 for distributed Jobs) |
Preempted, NodeLost |
recovery.onInterruption: Fail was set, so the first interruption ends the run |
Failures caused by your command (a non-zero exit, out of memory, an invalid checkpoint) fail fast and are not retried. At most three recoveries per organization prepare capacity at the same time; the others wait their turn and show an Event.
Suspending
Section titled “Suspending”nodus suspend job/<name> stops the run after a final checkpoint and releases its compute; nodus resume
continues at a new epoch from that checkpoint. If the final checkpoint fails or takes longer than
max(10 minutes, 2 × the shutdown reserve), the run keeps running with Suspended=False, reason=SuspendFailed
and a SuspendFailed Event, so a suspend never throws work away. Stops caused by money (credits, budgets, the
maximum cost) or by a lifetime limit do not wait: they stop at the funded edge with whatever checkpoint exists.
A stop always wins over recovery. If the machine is lost while a suspend, cancel or money stop is in progress, the
attempt ends Cancelled with the stop’s reason and nothing is restarted or billed again; nodus resume starts the
next epoch as usual.
Who pays for recovery
Section titled “Who pays for recovery”You pay for the machines your run uses, at the rate frozen when each was acquired, from the moment the provider starts billing through the confirmed deletion of the machine. The bill splits each machine’s time into segments:
| Segment | Covers |
|---|---|
Boot |
Start-up: boot, image pull, readiness, and for gangs the wait at the start barrier |
Restore |
Start-up of an attempt that resumes: a recovery, a resume after a suspend, and for gangs the surviving machines’ wait from the restart to the next epoch’s start |
Running |
Your command running |
Teardown |
Stop to confirmed deletion, plus the provider’s billing increment |
A recovery therefore costs you the replacement’s Boot or Restore time and the lost machine’s Teardown.
A machine the provider refuses, or that never appears, before it is ready is replaced with another one, up to six
tries per attempt; after that the attempt fails with ReadinessFailed and the run’s retry rule applies.
Nodus pays, and never charges you, for:
- a replacement released because the old machine came back;
- extra machines prepared to start faster (hedges) that did not win;
- machines that failed Nodus’s own checks after creation, and Nodus internal failures;
- for distributed Jobs, probe failures on a qualified network path and outages of the Nodus mesh.
nodus billing usage --group-by segment and the Cost tab show every segment of every attempt.