# Run jobs on your own machines

> Enroll your hosts into a pool, route jobs to them first, forecast demand and automate routine actions.

Source: https://nodus-platform-site.pages.dev/docs/guides/pools/
Build revision: 211ad9f836655b1c3a2668c4693e442471f28614

A pool is a group of your own machines that Nodus schedules work onto before it rents anything. Nodus does not charge a compute rental fee for your hosts; you continue paying your cloud or hardware costs. Routing a job to a pool costs $0.02 per GPU-hour while a Nodus-scheduled attempt uses a GPU. CPU-only pool work has no device-hour fee. Predict capacity forecasts cost $99.00 per pool per month. You agree to each price when you turn it on.

Pools have four parts:

* **Measure**: enroll hosts with a one-line installer and see their inventory, utilization and health.
* **Route**: send jobs to the pool first, wait for it or burst to the market under rules you set.
* **Predict**: forecasts of demand and free capacity, and recommendations for sizing and wait policies.
* **Act**: routine actions such as reclaiming idle hosts or draining in a maintenance window, with approvals.

## Create a pool

Terminal window

```sh
nodus create pool lab --routing prefer
```

The CLI and the console ask you to agree to the routing fee before the pool is created. In the console, open **BYOCompute** and choose **New pool**.

## Enroll a host

A host needs Linux on x86-64 or arm64 with a running systemd service manager and cgroup v2 with CPU and memory controllers. It also needs running containerd at `/run/containerd/containerd.sock`, a healthy `overlayfs` snapshotter, and `ctr`, `runc`, `containerd-shim-runc-v2` and `timeout` on its system path. Follow the [containerd setup instructions](https://github.com/containerd/containerd/blob/main/docs/getting-started.md) if these are missing. GPU hosts also need working NVIDIA drivers and the [NVIDIA Container Toolkit with CDI](https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/cdi-support.html).

The installer checks these prerequisites before downloading the agent or consuming the enrollment token. It does not install system packages, reconfigure Docker or restart containerd. Download the installer from the URL in the enrollment command to `install.sh`, then check the host without enrolling it:

Terminal window

```sh
sudo sh ./install.sh --check
```

A successful check confirms runtime prerequisites, not a completed GPU Job. Create an enrollment token for the pool; the response shows the installer once:

Terminal window

```sh
nodus create enrollmenttoken --pool lab
```

`--ttl 2h` shortens the token’s life and `--label rack=a` copies a label onto the node when it enrolls. The token needs no name; the server gives it one.

Run the command it prints as a user who can `sudo` on the host. It writes the token to a file only you can read, downloads the signed `nodusd` agent, checks its sha256 checksum, installs it as a systemd service and joins the pool. It also verifies the signature when `cosign` is installed; `--require-signature` makes that check mandatory. The token works once and expires after 24 hours; the console shows the countdown. Wait for the node to report `Ready` before submitting work. The installer is for enrollment; upgrading an existing host requires draining its work, replacing the verified agent binary and restarting `nodusd`.

Terminal window

```sh
nodus get nodes
nodus get node/host-1 -o yaml
```

Each job runs in a network namespace of its own, and `nodusd` keeps its network rules in the nftables table `inet nodus`. Outside that table it adds two rules to the host’s firewall, `-i ndv+ -j ACCEPT` and `-o ndv+ -j ACCEPT`. They match only its job interfaces, because Docker and ufw drop all forwarded traffic and jobs with open egress could not reach the internet otherwise. With Docker installed they go at the head of the `DOCKER-USER` chain; without Docker they go at the head of `FORWARD`, and only when its policy is `DROP`. Your rules for every other interface are untouched, and `inet nodus` still blocks private, link-local and metadata addresses for every job. `nodusd` puts the two rules back if they go missing and removes them when it stops.

A node reports its GPUs, CPU, memory and disk, its utilization and a heartbeat. A node with no heartbeat for 10 minutes shows **Offline** and the pool’s members get one email about it.

To stop new work on a host without stopping what runs there, drain it; undrain it to take work again. Deleting a node drains it, then revokes its credential:

Terminal window

```sh
nodus node drain host-1
nodus node undrain host-1
nodus delete node/host-1
```

## Route jobs to the pool

Name the pool in a job’s placement. With `mode: Prefer` the pool is used first and the market takes what the pool cannot; with `mode: Only` work never leaves the pool.

In the console, open your pool and choose **Run a job**. The job form selects your private pool under **Where to run**. Review the image, command and resources before launching. Your pool remains private to your organization.

Private execution currently supports single-node Jobs, including indexed jobs. Each attempt reserves one whole host, so another attempt waits until that host is released. A GPU attempt sees only the GPUs it requested and pays the routing fee for those GPUs. Distributed jobs and other resource kinds cannot use this route yet.

GPU jobs wait when the requested devices have existing GPU work, resident GPU memory, or no fresh utilization reading. After Nodus assigns a device, keep other host processes off it until the job ends; the host owner still controls processes started outside Nodus.

```yaml
apiVersion: nodus.dev/v1
kind: Job
metadata:
  name: train
spec:
  image: nodus/pytorch:2.8-cuda12.8
  command: [python, train.py]
  resources: {gpu: "H100:1"}
  placement: {pool: lab}
```

The pool’s `routing` decides what happens when it is full:

|Field|Values|Meaning|
|-|-|-|
|`waitPolicy`|`Never`, `Cheaper`, `Timeout`|Go to the market at once, wait while waiting costs less than the market, or wait up to `waitTimeout`.|
|`burstToMarket`|`Allow`, `Approve`, `Deny`|Burst at market rates, wait for an approval, or never burst.|

With `Approve`, a job that would burst shows `BurstApprovalRequired` until someone approves it:

Terminal window

```sh
nodus request approve-burst job/train
```

A job on the pool is billed only the routing fee for its device-hours. A burst to the market is billed like any other job.

## Forecast demand

Pool Predict is currently unavailable while the updated host telemetry is being qualified. Enabling it is rejected before charging. Cloud-spend projections in BYOCompute are available separately.

Turn on Predict on the pool’s **Forecast** tab or with `spec.predict.enabled: true`. Forecasts start once the pool has 7 days of utilization history, then refresh hourly; recommendations refresh daily.

Terminal window

```sh
nodus pool utilization lab
nodus pool forecast lab --horizon 7d
nodus pool recommendations lab
```

The forecast gives demand and free capacity in GPUs per hour with p50 and p90 bands. Recommendations cover right-sizing, idle hosts, maintenance windows, fragmentation and wait-policy tuning. Dismissing one records your reason, and the same recommendation is not raised again. Turning Predict off stops the charge at once.

Host utilization history requires an updated `nodusd` reporting GPU readings. Missing samples and long gaps do not establish idle capacity, and upgrading does not reconstruct past readings. The BYOCompute **Forecast** section also shows cloud-spend projections; those use your cloud billing history and are separate from Pool Predict capacity forecasts.

## Automate actions

`spec.act` sets the mode and the policies. Every action is a `PoolAction` you can list, and every executed action is audited.

|Mode|What happens|
|-|-|
|`Off`|Nothing is proposed.|
|`Shadow`|Records what would run and runs nothing.|
|`Propose`|Each action waits for approval and expires after 24 hours.|
|`Auto`|Runs actions within the policies at once.|

Terminal window

```sh
nodus get poolactions
nodus pool approve lab-idle-1
nodus pool reject lab-idle-1
nodus pool revert lab-idle-1
```

Actions only drain and undrain hosts in the pool, adjust its routing within the policy’s bounds, or approve a burst. They never stop or delete your jobs. To stop all actions at once, pause them; the pool shows `ActPaused` until you resume:

Terminal window

```sh
nodus pool pause lab
nodus pool resume lab
```

## See your cloud accounts

To list GPU instances and spend in your AWS or GCP accounts and enroll them into pools, connect them read-only: see [Connect your cloud accounts](https://nodus-platform-site.pages.dev/docs/guides/pools/cloud-accounts/).
