Placement and scheduling profiles
View MarkdownBy default Nodus places each run on the cheapest offering that can start it now: the lowest hourly rate among the offerings that fit your request, have a machine free and stay within your limits. If that offering cannot be used, the next cheapest takes the run, and so on. You pay the rate of the offering the run starts on, shown in the estimate and frozen for the run. This page explains what the estimate includes and the settings that change the choice.
Cost to completion
Section titled “Cost to completion”For every offering that fits your request, Nodus estimates the whole bill of the run. The estimate shows it for the
chosen offering, and the Cost profile ranks offerings on it:
- Startup: the machine’s boot, the image pull and any restore from a checkpoint. You pay for these because the capacity bills from the moment it is created.
- Running time: from your declared
expectedDuration, a training runtime’s measured estimate, or the history of earlier runs with the same labels on the same accelerator family. - Teardown: the time until the machine is confirmed deleted.
- Rounding: each machine bills in increments, and the estimate rounds up the same way.
- Interruptions: for interruptible capacity, the expected number of reclaims times the work redone after each one. Work that checkpoints loses only the stretch since its last save; work that saves nothing loses half its run on average.
Under Cost, a cheap interruptible offering therefore wins only when its expected cost, lost work included, still
beats on-demand capacity. A run whose checkpoints are all empty is costed as if it saves nothing.
When the running time is unknown, every profile ranks offerings by hourly rate and the estimate shows the rate and the startup and teardown charge but no total.
Profiles
Section titled “Profiles”placement.profile picks how the scheduler trades cost against time:
| Profile | Chooses |
|---|---|
Balanced (default) |
The lowest hourly rate that meets your deadline and budget. Among offerings within 2 % of that rate, the healthier one, or the one that starts sooner, fits better or is already warm |
Cost |
The lowest expected cost to completion, startup, teardown and lost work included, even if it is slower to start. Among offerings within 2 % of that cost, warm capacity and a better fit win |
Speed |
The fastest expected finish among offerings within 1.5 × the cheapest expected cost |
Balanced never chooses an offering whose rate is more than 2 % above the cheapest that meets your limits, and
Cost never chooses one whose expected cost is more than 2 % above the cheapest at equal health;
interruptible: Prefer gives interruptible capacity a 10 % allowance. Under Cost and Speed, an offering that
often fails to start counts as dearer by that risk; under every profile, one that recently failed to start ranks
lower for a few minutes within the 2 % band. When two offerings are within 2 % of each other, identical requests are
spread across both instead of all taking the same one.
Deadlines and budgets
Section titled “Deadlines and budgets”placement.completeByTime: offerings whose p90 finish is later are not used.maxCostUSD, Budgets and your balance: offerings whose expected cost exceeds the money left are not used.placement.maxRateUSDPerHour: offerings above this rate are not used.
If nothing remains, the run waits in Queued and the estimate says why, for example MissesDeadline or
ExceedsRemainingBudget.
Estimates and the If-Match ceiling
Section titled “Estimates and the If-Match ceiling”--dry-run=server -o estimate returns the expected cost p50 and p90, the startup time (cold, and warm when idle
capacity of yours fits), the first hold, the minimum charge and validUntil, which is at most 31 minutes away.
Creating with the estimate’s If-Match binds the launch to it: Nodus then uses no offering above the estimated
rate plus 10 %, unless you set placement.maxRateUSDPerHour yourself. If prices moved beyond that, the run waits
with PriceAboveEstimate instead of costing more than you saw.
Reuse before new capacity
Section titled “Reuse before new capacity”Before buying new capacity, the scheduler considers capacity you already pay for: an idle machine of yours that
fits reuses the increment already paid, and CPU work packs onto your Nodus nodes. A BYOC pool
named in placement.pool is used first at no hourly charge.
Why a placement was made
Section titled “Why a placement was made”nodus describe shows each attempt’s placement: the profile, the scores of the chosen offering (fit, time to
result, cost to complete, recovery value, health), the fallbacks in order and every rejected offering with its
reason. Offerings are shown by name, such as h100-sxm-80g-x8-us.
Multi-node runs (Beta)
Section titled “Multi-node runs (Beta)”For spec.distributed, Nodus first resolves the topology (nodes × GPUs per node) and then places every node
together:
network: Colocated(the default) keeps every node in one location on one network;Regionalkeeps them in one region class;Globalallows anywhere.transport: Directuses only private or direct paths between nodes;Autoalso allows the relayed mesh, with lower bandwidth.- The whole gang is priced together, from the slowest node’s start. The estimate also shows the assembly bound: the most a failed assembly can cost.
If no set of offerings satisfies these rules, the run waits with GangInfeasible. See
multi-node training.