# Placement and scheduling profiles

> How Nodus chooses where a run executes, what the estimate includes, and how profiles, deadlines and budgets change the choice.

Source: https://nodus-platform-site.pages.dev/docs/concepts/supply-placement/
Build revision: 211ad9f836655b1c3a2668c4693e442471f28614

By default Nodus places each run on the **cheapest offering that can start it now**: the lowest hourly rate among the offerings that fit your request, have a machine free and stay within your limits. If that offering cannot be used, the next cheapest takes the run, and so on. You pay the rate of the offering the run starts on, shown in the estimate and frozen for the run. This page explains what the estimate includes and the settings that change the choice.

## Cost to completion

For every offering that fits your request, Nodus estimates the whole bill of the run. The estimate shows it for the chosen offering, and the `Cost` profile ranks offerings on it:

* **Startup**: the machine’s boot, the image pull and any restore from a checkpoint. You pay for these because the capacity bills from the moment it is created.
* **Running time**: from your declared `expectedDuration`, a training runtime’s measured estimate, or the history of earlier runs with the same labels on the same accelerator family.
* **Teardown**: the time until the machine is confirmed deleted.
* **Rounding**: each machine bills in increments, and the estimate rounds up the same way.
* **Interruptions**: for interruptible capacity, the expected number of reclaims times the work redone after each one. Work that checkpoints loses only the stretch since its last save; work that saves nothing loses half its run on average.

Under `Cost`, a cheap interruptible offering therefore wins only when its expected cost, lost work included, still beats on-demand capacity. A run whose checkpoints are all empty is costed as if it saves nothing.

When the running time is unknown, every profile ranks offerings by hourly rate and the estimate shows the rate and the startup and teardown charge but no total.

## Profiles

`placement.profile` picks how the scheduler trades cost against time:

|Profile|Chooses|
|-|-|
|`Balanced` (default)|The lowest hourly rate that meets your deadline and budget. Among offerings within 2 % of that rate, the healthier one, or the one that starts sooner, fits better or is already warm|
|`Cost`|The lowest expected cost to completion, startup, teardown and lost work included, even if it is slower to start. Among offerings within 2 % of that cost, warm capacity and a better fit win|
|`Speed`|The fastest expected finish among offerings within 1.5 × the cheapest expected cost|

`Balanced` never chooses an offering whose rate is more than 2 % above the cheapest that meets your limits, and `Cost` never chooses one whose expected cost is more than 2 % above the cheapest at equal health; `interruptible: Prefer` gives interruptible capacity a 10 % allowance. Under `Cost` and `Speed`, an offering that often fails to start counts as dearer by that risk; under every profile, one that recently failed to start ranks lower for a few minutes within the 2 % band. When two offerings are within 2 % of each other, identical requests are spread across both instead of all taking the same one.

## Deadlines and budgets

* `placement.completeByTime`: offerings whose p90 finish is later are not used.
* `maxCostUSD`, Budgets and your balance: offerings whose expected cost exceeds the money left are not used.
* `placement.maxRateUSDPerHour`: offerings above this rate are not used.

If nothing remains, the run waits in `Queued` and the estimate says why, for example `MissesDeadline` or `ExceedsRemainingBudget`.

## Estimates and the If-Match ceiling

`--dry-run=server -o estimate` returns the expected cost p50 and p90, the startup time (cold, and warm when idle capacity of yours fits), the first hold, the minimum charge and `validUntil`, which is at most 31 minutes away. Creating with the estimate’s `If-Match` binds the launch to it: Nodus then uses no offering above the estimated rate plus 10 %, unless you set `placement.maxRateUSDPerHour` yourself. If prices moved beyond that, the run waits with `PriceAboveEstimate` instead of costing more than you saw.

## Reuse before new capacity

Before buying new capacity, the scheduler considers capacity you already pay for: an idle machine of yours that fits reuses the increment already paid, and CPU work packs onto your Nodus nodes. A [BYOC pool](https://nodus-platform-site.pages.dev/docs/guides/pools/) named in `placement.pool` is used first at no hourly charge.

## Why a placement was made

`nodus describe` shows each attempt’s placement: the profile, the scores of the chosen offering (fit, time to result, cost to complete, recovery value, health), the fallbacks in order and every rejected offering with its reason. Offerings are shown by name, such as `h100-sxm-80g-x8-us`.

## Multi-node runs (Beta)

For `spec.distributed`, Nodus first resolves the topology (nodes × GPUs per node) and then places every node together:

* `network: Colocated` (the default) keeps every node in one location on one network; `Regional` keeps them in one region class; `Global` allows anywhere.
* `transport: Direct` uses only private or direct paths between nodes; `Auto` also allows the relayed mesh, with lower bandwidth.
* The whole gang is priced together, from the slowest node’s start. The estimate also shows the assembly bound: the most a failed assembly can cost.

If no set of offerings satisfies these rules, the run waits with `GangInfeasible`. See [multi-node training](https://nodus-platform-site.pages.dev/docs/guides/multi-node/).
