# Choose GPUs and check availability

> See which accelerators are available right now, what they cost, and how to ask for exactly the hardware your run needs.

Source: https://nodus-platform-site.pages.dev/docs/guides/gpus/
Build revision: 211ad9f836655b1c3a2668c4693e442471f28614

You describe the hardware a run needs; Nodus finds capacity that fits and finishes it for the lowest expected cost. This guide shows how to see what is available, read an offering, and write a GPU request that says exactly what you mean.

## See what is available

Terminal window

```console
$ nodus get gpus
NAME                  TYPE       COUNT   VCPU   MEMORY   REGION    PRICE/H   AVAILABILITY   AGE
a100-sxm-80g-x1-any   A100-80G   1       30     200Gi    unknown   $1.27     Available      14s
h100-sxm-80g-x8-us    H100-SXM   8       208    1800Gi   us        $18.29    Available      14s
l4-24g-x1-eu-int      L4         1       8      32Gi     eu        $0.31     Limited        14s
```

`PRICE/H` is the from rate per machine-hour and `AGE` is how long ago the price and availability were confirmed.

Filter with flags; each one narrows the list:

Terminal window

```console
$ nodus get gpus --gpu H100 --count 8 --region us
$ nodus get gpus --interruptible
$ nodus get gpus --gpu L4 -o wide        # typical and list rates, startup time and interruption rate
```

The console’s GPU picker shows the same list.

## Read an offering

An offering is a class of capacity: accelerator × count × region class, and whether it is interruptible. Its name spells that out:

|Name|Meaning|
|-|-|
|`h100-sxm-80g-x8-us`|Eight H100 SXM 80 GB GPUs per machine, in the `us` region class|
|`l4-24g-x1-eu-int`|One L4 per machine in `eu`, interruptible|
|`a100-sxm-80g-x1-any`|Capacity with no guaranteed location (`any`)|
|`cpu-8c-32g-us`|An 8 vCPU, 32 GiB CPU machine in `us`|

Each offering shows:

* **From** and **typical** rates per hour. The from rate is that of the cheapest machine free now, the one a run starts on by default, or of the cheapest machine when none is free; the host shape shown is that machine’s. Your rate is shown in the estimate before launch and frozen for the run. No rate is ever above the list price.
* **Availability**: `Available`, `Limited` (a few machines, or a count that is not reported) or `Unavailable`.
* **Startup**: how long a machine usually takes to be ready (p50 and p90), from what Nodus has measured.
* **Interruption rate** for interruptible offerings: how often such capacity is taken back per hour.

Prices and availability refresh continuously. The public [pricing page](https://nodus-platform-site.pages.dev/pricing) shows the same from rates.

## Region classes

`placement.regions` takes region classes, not data centers: `us`, `ca`, `eu`, `uk`, `in`, `apac`, `me` and `latam`. Capacity whose location is not reported has the class `any` in its name and never satisfies a non-empty `placement.regions`, so a run that must stay in a region never lands there.

```yaml
spec:
  placement:
    regions: [eu, uk]
```

## Ask for the hardware you need

`resources.gpu.type` lists every accelerator you accept. Order does not matter: the scheduler picks among them by expected cost to finish.

|You write|Nodus may use|
|-|-|
|`H100`|Any H100 variant: SXM, PCIe or NVL|
|`[H100, H200]`|Any H100 or H200, never another family|
|`H100-SXM`|Only the H100 SXM|
|`H100!` or `exact: true`|Only the family’s primary variant (H100 SXM)|
|`minMemory: 80Gi` with no type|Any accelerator with at least 80 GiB per GPU|

Modal spellings work as written: `A100-40GB`, `A100-80GB`, `A10G`, `L40S`, `H100`, `H200`, `B200`, `T4`, `L4`.

```yaml
spec:
  resources:
    gpu: {type: [H100, H200], count: 8}
  placement:
    interruptible: Allow          # use interruptible capacity only when it is cheaper to finish
    maxRateUSDPerHour: "30.00"
```

`count` is per machine (1, 2, 4 or 8). For more GPUs than one machine holds, see [multi-node training](https://nodus-platform-site.pages.dev/docs/guides/multi-node/) (Beta).

## Use every GPU of one machine

`--gpu A6000:2` (or `count: 2`) gives your container both GPUs of one machine, numbered 0 and 1. `torchrun` starts one process per GPU without flags, because Nodus sets `PET_NPROC_PER_NODE` to the GPU count; a `--nproc-per-node` you pass yourself wins.

train.py

```python
"""A DDP smoke run on every GPU of one machine. torchrun starts one process per GPU (PET_NPROC_PER_NODE)."""
import os

import torch
import torch.distributed as dist
from torch.nn.parallel import DistributedDataParallel as DDP

dist.init_process_group("nccl")
local_rank = int(os.environ["LOCAL_RANK"])
device = torch.device("cuda", local_rank)
torch.cuda.set_device(device)

# Every process contributes 1, so the sum is the world size.
one = torch.ones(1, device=device)
dist.all_reduce(one)

model = DDP(torch.nn.Linear(16, 1).to(device), device_ids=[local_rank])
opt = torch.optim.SGD(model.parameters(), lr=0.1)
for step in range(20):
    x = torch.randn(32, 16, device=device)
    loss = (model(x) - x.sum(dim=1, keepdim=True)).pow(2).mean()
    opt.zero_grad()
    loss.backward()
    opt.step()

if dist.get_rank() == 0:
    print(f"world_size={int(one.item())} gpus={torch.cuda.device_count()} "
          f"device={torch.cuda.get_device_name(device)} loss={loss.item():.4f}")
dist.destroy_process_group()
```

run.sh

```sh
nodus run --name torchrun-multi-gpu --gpu A6000:2 --max-cost 1.00 -- torchrun train.py
```

The run prints `world_size=2 gpus=2` from the first process. The processes reach each other over NCCL on the machine itself; `/dev/shm` is sized for it (half the container’s memory), so you need no `--shm-size` or `--ipc=host`.

## Check before you launch

A dry run returns the estimate without starting anything:

Terminal window

```console
$ nodus apply -f job.yaml --dry-run=server -o estimate
```

It shows the offering Nodus would use, the expected cost p50 and p90, the startup time, the first hold and the minimum charge, and how long the estimate stays valid. When nothing fits, it says why, for example `CapacityUnavailable`, `NoListPrice`, `ExceedsRemainingBudget` or `MissesDeadline`, with a fix.

## When a run waits in Queued

A run with no matching capacity stays `Queued` and is placed as soon as capacity appears, until `placement.queueTimeout` (24 hours by default). `nodus describe` shows the offerings that were considered and why each was rejected. To start sooner, widen `resources.gpu.type` or `placement.regions`, allow interruptible capacity, or raise `placement.maxRateUSDPerHour`.

How Nodus chooses among offerings is explained in [placement and scheduling profiles](https://nodus-platform-site.pages.dev/docs/concepts/supply-placement/).
