# Sweeps

> Run one Job across GPU types, regions and parameter values, and compare cost, time and throughput per cell.

Source: https://nodus-platform-site.pages.dev/docs/guides/sweeps/
Build revision: 211ad9f836655b1c3a2668c4693e442471f28614

A Sweep runs the same Job once for every combination of the values you list: GPU types, region classes and your own parameters. Each combination is a cell. When the cells finish, the Sweep shows what each one cost, how long it took and how fast it went, and points out the cheapest and the fastest.

## Submit a Sweep

This Sweep measures training throughput for three batch sizes on two GPU types, six cells in all:

sweep.yaml

```yaml
apiVersion: nodus.dev/v1
kind: Sweep
metadata:
  name: batch-size
spec:
  maxCostUSD: "0.75"
  maxParallel: 2
  # 2 GPU types x 3 batch sizes = 6 cells, each a Job named batch-size-<index>.
  matrix:
    gpu: [L4, A10]
    params:
      BATCH_SIZE: ["32", "64", "128"]
  template:
    kind: Job
    spec:
      image: nodus/pytorch
      timeout: 10m
      command:
        - python
        - -c
        - |
          import os, time, torch
          bs = int(os.environ["NODUS_PARAM_BATCH_SIZE"])
          model = torch.nn.Sequential(torch.nn.Linear(1024, 4096), torch.nn.ReLU(), torch.nn.Linear(4096, 10)).cuda()
          opt = torch.optim.AdamW(model.parameters())
          x, y = torch.randn(bs, 1024).cuda(), torch.randint(0, 10, (bs,)).cuda()
          start, steps = time.time(), 300
          for _ in range(steps):
              opt.zero_grad()
              torch.nn.functional.cross_entropy(model(x), y).backward()
              opt.step()
          torch.cuda.synchronize()
          print(f"batch {bs}: {steps * bs / (time.time() - start):.0f} samples/s")
```

Terminal window

```console
$ nodus apply -f sweep.yaml
sweep.nodus.dev/batch-size created
```

Each parameter reaches the command as an environment variable: `BATCH_SIZE` arrives as `NODUS_PARAM_BATCH_SIZE`. Parameter names use uppercase letters, digits and `_`.

## Watch it and read the results

Terminal window

```console
$ nodus get sweep/batch-size -w
NAME         PHASE       CELLS   COST    BEST-COST   AGE
batch-size   Succeeded   6/6     $0.31   2           14m
$ nodus get sweep/batch-size -o yaml
```

Each cell runs as a Job named `<sweep>-<index>`, so `nodus logs job/batch-size-3` shows one cell’s output. The Sweep’s `status.cells` lists every cell with its GPU, region, parameters, phase, `costUSD`, `wallSeconds` and `unitsPerSecond`, and `status.best` names the cell with the lowest cost (`byCost`) and the highest throughput (`byThroughput`).

`unitsPerSecond` counts the units your program reports with `nodus.log.unit(id, ms)` from the Python SDK, such as one call per batch or per request, divided by the cell’s wall time. Cells that report no units have no throughput.

## The matrix

|Field|Values|Limit|
|-|-|-|
|`matrix.gpu`|GPU requests, such as `L4`, `H100:8` or `H100!`|16|
|`matrix.regions`|Region classes, such as `us` or `eu`; within the template’s `placement.regions` when it sets them||
|`matrix.params`|Parameter name → list of string values|16 names, 64 values each|
|`repetitions`|Runs of each combination, to measure variance|64|

The number of cells is the product of the non-empty dimensions times `repetitions`, at most 256. Cells are numbered in a fixed order: GPU types vary slowest, then regions, then parameters by name, and repetitions fastest, so the repetitions of one combination sit next to each other.

## From Python

```python
import nodus

sweep = nodus.Sweep(
    {"image": "nodus/pytorch", "command": ["python", "bench.py"]},     # the Job spec every cell runs
    grid={"gpu": ["L4", "H100"], "BATCH_SIZE": [8, 16, 32]},           # gpu and region are dimensions; the rest are parameters
    repetitions=2, max_parallel=6, max_cost=25,
).run()
report = sweep.wait()                  # a Sweep with failed cells is Failed (CellsFailed) and still has its report
print(report.phase, report.best.by_cost)
for cell in sweep.cells():
    print(cell.index, cell.phase, cell.cost_usd, cell.wall_seconds)
```

`nodus.Sweep.from_name("batch-size")` reads one that exists, and `suspend()`, `resume()` and `cancel()` apply to every unfinished cell. A Function or a recipe as the target is not available yet and raises `nodus.errors.Unsupported`.

## Running and failing cells

* `maxParallel` limits how many cells run at once (default 4). Cells start in index order.
* `maxCostUSD` caps the whole Sweep. At the cap the running cells save their state and suspend, no new cell starts, and the Sweep becomes `Suspended` with reason `MaxCostReached`; raise the cap to continue.
* A failed cell does not stop the others. When every cell has finished and some failed, the Sweep fails with reason `CellsFailed`, and its `status.cells` still carries the results of the cells that succeeded.
* `nodus suspend sweep/x`, `resume` and `cancel` apply to every unfinished cell. Deleting a Sweep deletes its cell Jobs.
