Skip to content

Run jobs on your own machines

View Markdown

A pool is a group of your own machines that Nodus schedules work onto before it rents anything. Nodus does not charge a compute rental fee for your hosts; you continue paying your cloud or hardware costs. Routing a job to a pool costs $0.02 per GPU-hour while a Nodus-scheduled attempt uses a GPU. CPU-only pool work has no device-hour fee. Predict capacity forecasts cost $99.00 per pool per month. You agree to each price when you turn it on.

Pools have four parts:

  • Measure: enroll hosts with a one-line installer and see their inventory, utilization and health.
  • Route: send jobs to the pool first, wait for it or burst to the market under rules you set.
  • Predict: forecasts of demand and free capacity, and recommendations for sizing and wait policies.
  • Act: routine actions such as reclaiming idle hosts or draining in a maintenance window, with approvals.
Terminal window
nodus create pool lab --routing prefer

The CLI and the console ask you to agree to the routing fee before the pool is created. In the console, open BYOCompute and choose New pool.

A host needs Linux on x86-64 or arm64 with a running systemd service manager and cgroup v2 with CPU and memory controllers. It also needs running containerd at /run/containerd/containerd.sock, a healthy overlayfs snapshotter, and ctr, runc, containerd-shim-runc-v2 and timeout on its system path. Follow the containerd setup instructions if these are missing. GPU hosts also need working NVIDIA drivers and the NVIDIA Container Toolkit with CDI.

The installer checks these prerequisites before downloading the agent or consuming the enrollment token. It does not install system packages, reconfigure Docker or restart containerd. Download the installer from the URL in the enrollment command to install.sh, then check the host without enrolling it:

Terminal window
sudo sh ./install.sh --check

A successful check confirms runtime prerequisites, not a completed GPU Job. Create an enrollment token for the pool; the response shows the installer once:

Terminal window
nodus create enrollmenttoken --pool lab

--ttl 2h shortens the token’s life and --label rack=a copies a label onto the node when it enrolls. The token needs no name; the server gives it one.

Run the command it prints as a user who can sudo on the host. It writes the token to a file only you can read, downloads the signed nodusd agent, checks its sha256 checksum, installs it as a systemd service and joins the pool. It also verifies the signature when cosign is installed; --require-signature makes that check mandatory. The token works once and expires after 24 hours; the console shows the countdown. Wait for the node to report Ready before submitting work. The installer is for enrollment; upgrading an existing host requires draining its work, replacing the verified agent binary and restarting nodusd.

Terminal window
nodus get nodes
nodus get node/host-1 -o yaml

Each job runs in a network namespace of its own, and nodusd keeps its network rules in the nftables table inet nodus. Outside that table it adds two rules to the host’s firewall, -i ndv+ -j ACCEPT and -o ndv+ -j ACCEPT. They match only its job interfaces, because Docker and ufw drop all forwarded traffic and jobs with open egress could not reach the internet otherwise. With Docker installed they go at the head of the DOCKER-USER chain; without Docker they go at the head of FORWARD, and only when its policy is DROP. Your rules for every other interface are untouched, and inet nodus still blocks private, link-local and metadata addresses for every job. nodusd puts the two rules back if they go missing and removes them when it stops.

A node reports its GPUs, CPU, memory and disk, its utilization and a heartbeat. A node with no heartbeat for 10 minutes shows Offline and the pool’s members get one email about it.

To stop new work on a host without stopping what runs there, drain it; undrain it to take work again. Deleting a node drains it, then revokes its credential:

Terminal window
nodus node drain host-1
nodus node undrain host-1
nodus delete node/host-1

Name the pool in a job’s placement. With mode: Prefer the pool is used first and the market takes what the pool cannot; with mode: Only work never leaves the pool.

In the console, open your pool and choose Run a job. The job form selects your private pool under Where to run. Review the image, command and resources before launching. Your pool remains private to your organization.

Private execution currently supports single-node Jobs, including indexed jobs. Each attempt reserves one whole host, so another attempt waits until that host is released. A GPU attempt sees only the GPUs it requested and pays the routing fee for those GPUs. Distributed jobs and other resource kinds cannot use this route yet.

GPU jobs wait when the requested devices have existing GPU work, resident GPU memory, or no fresh utilization reading. After Nodus assigns a device, keep other host processes off it until the job ends; the host owner still controls processes started outside Nodus.

apiVersion: nodus.dev/v1
kind: Job
metadata:
name: train
spec:
image: nodus/pytorch:2.8-cuda12.8
command: [python, train.py]
resources: {gpu: "H100:1"}
placement: {pool: lab}

The pool’s routing decides what happens when it is full:

Field Values Meaning
waitPolicy Never, Cheaper, Timeout Go to the market at once, wait while waiting costs less than the market, or wait up to waitTimeout.
burstToMarket Allow, Approve, Deny Burst at market rates, wait for an approval, or never burst.

With Approve, a job that would burst shows BurstApprovalRequired until someone approves it:

Terminal window
nodus request approve-burst job/train

A job on the pool is billed only the routing fee for its device-hours. A burst to the market is billed like any other job.

Pool Predict is currently unavailable while the updated host telemetry is being qualified. Enabling it is rejected before charging. Cloud-spend projections in BYOCompute are available separately.

Turn on Predict on the pool’s Forecast tab or with spec.predict.enabled: true. Forecasts start once the pool has 7 days of utilization history, then refresh hourly; recommendations refresh daily.

Terminal window
nodus pool utilization lab
nodus pool forecast lab --horizon 7d
nodus pool recommendations lab

The forecast gives demand and free capacity in GPUs per hour with p50 and p90 bands. Recommendations cover right-sizing, idle hosts, maintenance windows, fragmentation and wait-policy tuning. Dismissing one records your reason, and the same recommendation is not raised again. Turning Predict off stops the charge at once.

Host utilization history requires an updated nodusd reporting GPU readings. Missing samples and long gaps do not establish idle capacity, and upgrading does not reconstruct past readings. The BYOCompute Forecast section also shows cloud-spend projections; those use your cloud billing history and are separate from Pool Predict capacity forecasts.

spec.act sets the mode and the policies. Every action is a PoolAction you can list, and every executed action is audited.

Mode What happens
Off Nothing is proposed.
Shadow Records what would run and runs nothing.
Propose Each action waits for approval and expires after 24 hours.
Auto Runs actions within the policies at once.
Terminal window
nodus get poolactions
nodus pool approve lab-idle-1
nodus pool reject lab-idle-1
nodus pool revert lab-idle-1

Actions only drain and undrain hosts in the pool, adjust its routing within the policy’s bounds, or approve a burst. They never stop or delete your jobs. To stop all actions at once, pause them; the pool shows ActPaused until you resume:

Terminal window
nodus pool pause lab
nodus pool resume lab

To list GPU instances and spend in your AWS or GCP accounts and enroll them into pools, connect them read-only: see Connect your cloud accounts.