Skip to content

Attempts and recovery

View Markdown

Every run on Nodus executes as one or more attempts. An attempt is one incarnation of your command on one machine. When that machine is reclaimed, loses its network or fails, Nodus starts a new attempt somewhere else and your run continues from what it saved. This page explains what carries over, how Nodus decides a machine is gone, and what you pay for along the way.

A Job index, a Sandbox or a worker slot holds one running attempt at a time. Each new attempt gets the next epoch, a number that only grows. Nodus accepts reports, checkpoints and outputs only from the current epoch, so a machine that comes back after it was replaced can never overwrite the work of its successor.

You see a run’s attempts with nodus get attempts -l nodus.dev/job=<name>. A run’s attempts share its name with an epoch suffix, and gang members add a rank suffix (-r1, -r2).

Attempt phase Meaning
Pending, Placing, Acquiring Choosing and preparing capacity
Starting The machine is ready; the image is pulled and inputs or a checkpoint are restored
Running Your command is running
Succeeded Your command exited 0
Failed Your command or its machine failed; reason says which (NodeLost, Preempted, OOMKilled, …)
Cancelled The attempt was stopped on purpose: a suspend, a cancel, a budget or lifetime limit

recovery.continuity says what a new attempt starts from:

Mode A new attempt starts with Use it for
Checkpointed The latest committed checkpoint of your declared state paths (NODUS_CHECKPOINT_DIR by default) Training and long jobs that save their own model, optimizer and progress files
Restartable A cold start plus the progress cursor you reported (NODUS_CURSOR_COMPLETED, NODUS_CURSOR_TOTAL) Batch work that can skip what it already finished
Ephemeral A cold start Short or idempotent work
Snapshotted The latest filesystem snapshot Sandboxes

Restoring files restores files only, never process memory. Your program loads its own checkpoint when it starts; Nodus never adds resume flags to your command.

An empty checkpoint never counts as saved progress and never replaces an earlier useful one. If every checkpoint a run commits is empty, the run shows Checkpointed=False, reason=NotCheckpointable, and Nodus plans and prices it as Ephemeral from then on.

When a provider reclaims interruptible capacity or a machine stops answering, Nodus recovers make-before-break:

  1. On a reclaim notice, Nodus asks your attempt for an urgent checkpoint and, at the same time, starts preparing a replacement machine.
  2. The replacement is prepared up to the point where it could start, but it does not start while the old machine might still be writing.
  3. It starts only once the old machine is provably gone: the old attempt acknowledged its stop, the provider confirmed the machine terminated, or the old attempt’s lease ran out.
  4. The new attempt restores according to the continuity mode above.

So a Checkpointed run loses at most the work since its last committed checkpoint, and a checkpoint committed during the reclaim notice still counts. If the old machine comes back before it is replaced, the replacement is released and your run simply continues; you are not charged for the replacement.

A machine that stops sending heartbeats gets a replacement prepared after 20 seconds. After 60 seconds of silence an attempt on shared capacity is declared lost and replaced. An attempt on a machine dedicated to your organization may keep running up to the edge of its funding while its replacement waits, because such a machine can only lose its own work.

Recovery stops, and the run fails with a typed reason, when:

Reason Rule
NoProgress Two attempts in a row made no progress. Progress means the command started and either ran for recovery.minProgressDuration (default 2 minutes) or committed a non-empty checkpoint. A preemption after progress never counts
RestoreFailed The same checkpoint failed to restore twice
ImagePullFailed The image failed to pull on a second machine (a pull failure is retried once elsewhere)
MaxAttemptsExceeded recovery.maxAttempts recoveries were used (default 8; 3 for distributed Jobs)
Preempted, NodeLost recovery.onInterruption: Fail was set, so the first interruption ends the run

Failures caused by your command (a non-zero exit, out of memory, an invalid checkpoint) fail fast and are not retried. At most three recoveries per organization prepare capacity at the same time; the others wait their turn and show an Event.

nodus suspend job/<name> stops the run after a final checkpoint and releases its compute; nodus resume continues at a new epoch from that checkpoint. If the final checkpoint fails or takes longer than max(10 minutes, 2 × the shutdown reserve), the run keeps running with Suspended=False, reason=SuspendFailed and a SuspendFailed Event, so a suspend never throws work away. Stops caused by money (credits, budgets, the maximum cost) or by a lifetime limit do not wait: they stop at the funded edge with whatever checkpoint exists.

A stop always wins over recovery. If the machine is lost while a suspend, cancel or money stop is in progress, the attempt ends Cancelled with the stop’s reason and nothing is restarted or billed again; nodus resume starts the next epoch as usual.

You pay for the machines your run uses, at the rate frozen when each was acquired, from the moment the provider starts billing through the confirmed deletion of the machine. The bill splits each machine’s time into segments:

Segment Covers
Boot Start-up: boot, image pull, readiness, and for gangs the wait at the start barrier
Restore Start-up of an attempt that resumes: a recovery, a resume after a suspend, and for gangs the surviving machines’ wait from the restart to the next epoch’s start
Running Your command running
Teardown Stop to confirmed deletion, plus the provider’s billing increment

A recovery therefore costs you the replacement’s Boot or Restore time and the lost machine’s Teardown. A machine the provider refuses, or that never appears, before it is ready is replaced with another one, up to six tries per attempt; after that the attempt fails with ReadinessFailed and the run’s retry rule applies. Nodus pays, and never charges you, for:

  • a replacement released because the old machine came back;
  • extra machines prepared to start faster (hedges) that did not win;
  • machines that failed Nodus’s own checks after creation, and Nodus internal failures;
  • for distributed Jobs, probe failures on a qualified network path and outages of the Nodus mesh.

nodus billing usage --group-by segment and the Cost tab show every segment of every attempt.