# Gang networking

> How the members of a distributed Job reach each other across machines and providers, how Nodus checks the path before training starts, and what each path can carry.

Source: https://nodus-platform-site.pages.dev/docs/concepts/gang-networking/
Build revision: 211ad9f836655b1c3a2668c4693e442471f28614

**Beta:** this feature may change.

Beta

Distributed Jobs (`Job.spec.distributed`) are in beta behind per-org access. Gangs have 2 to 8 members, and `Relayed` gangs have 2 until multi-member Relayed runs are qualified. Ask for access from the console.

A distributed Job runs as a **gang**: one member per machine, every member started together, sometimes on machines from different providers. Before your command starts, Nodus joins the members into a private network that belongs to that gang alone. Each rank gets the addresses of the others in its environment, so `torchrun`, Ray and plain `torch.distributed` work without any networking code of yours.

You choose two things in `spec.distributed`:

* `network`: how close the members must be. `Colocated` keeps them in one provider region, `Regional` allows any provider inside one region class, and `Global` allows anywhere.
* `transport`: which paths you accept. `Direct` (the default) accepts only `Private` and `Direct` paths. `Auto` also accepts `Relayed` paths, which lets members on machines without a kernel network device join.

Nodus picks the path. Every member of a gang uses the same path class, so every rank sees the same addressing and the same NCCL settings.

## Path classes

|Path|When Nodus uses it|How traffic flows|Encryption|Bandwidth|
|-|-|-|-|-|
|`Private`|Every member is in one provider region with a private network|The provider’s private network, opened only between the gang’s members|None added (the provider’s network)|Provider native|
|`Direct`|Every member can create a WireGuard device and the probe finds a direct path between every pair|WireGuard between the members, on a network device named `nodus0`|WireGuard|Measured per pair (see [measured throughput](https://nodus-platform-site.pages.dev/docs/concepts/gang-networking/#measured-throughput))|
|`Relayed`|A member has no network device of its own (container-only offerings) and `transport: Auto`|WireGuard in user space, carried through Nodus relays|WireGuard end to end; relays see only ciphertext|**Low: small models only**|

### Private

Members in the same provider region talk over the provider’s private network. Nodus opens the rendezvous and NCCL ports between the gang’s members only and closes them when the gang ends. The advertised addresses are the members’ private IPs.

### Direct

Each member runs a WireGuard endpoint next to your container, never inside it: your container gets no extra privileges. The members find each other through the Nodus mesh control and connect peer to peer, punching through NAT where needed. The advertised addresses are the members’ mesh addresses (`100.64.0.0/10`), and NCCL and Gloo use the `nodus0` interface.

`Direct` means direct. If the probe finds that one pair can only connect through a relay, a `transport: Direct` gang is placed again without that pair, and a `transport: Auto` gang continues as `Relayed`. If a running pair loses its direct path, its traffic falls back to a relay without interrupting training, and the Job’s `NetworkDegraded` condition becomes true.

### Relayed

Container-only offerings give your container no network device, no `CAP_NET_ADMIN` and no UDP. Their members can still join a gang: each runs WireGuard in user space and sends peer traffic through two relays that Nodus operates. To reach that WireGuard endpoint, your processes start with a small preloaded library, `libnodus-netshim.so`, which redirects only connections to the other members of your gang. Everything else, such as downloads, datasets, object storage and model APIs, goes out directly as usual.

The advertised addresses are the members’ own `eth0` addresses, so `MASTER_ADDR`, the `torchrun` rendezvous and the listeners NCCL opens on ephemeral ports all work unchanged, in both directions. The relays mesh with each other: if one goes down, members move to the other within seconds and the TCP connections inside WireGuard survive.

## What every rank sees

The path decides a few variables; the rest of the distributed environment is the same on every path.

|Variable|`Private`|`Direct`|`Relayed`|
|-|-|-|-|
|`NODUS_GANG_TRANSPORT`|`private`|`direct`|`relayed`|
|`MASTER_ADDR`, `PET_RDZV_ENDPOINT`, `NODUS_NODE_IPS`|Private IPs|Mesh addresses|`eth0` addresses|
|`NCCL_SOCKET_IFNAME`, `GLOO_SOCKET_IFNAME`|The private interface|`nodus0`|`eth0`|
|`LD_PRELOAD`, `NODUS_NETSHIM_PEERS`, `NODUS_NETSHIM_SOCKS`|Not set|Not set|The shim first, then your image’s own `LD_PRELOAD`|

Your own `NCCL_*` values, such as `NCCL_DEBUG=INFO`, win, except `NCCL_NET`, `NCCL_SOCKET_IFNAME`, `NCCL_SOCKET_FAMILY` and `NCCL_IB_DISABLE`, which the path fixes. A spec `LD_PRELOAD` is rejected when `transport: Auto`, because Nodus needs to place the shim first on a `Relayed` path. An `LD_PRELOAD` set in the image, such as jemalloc or tcmalloc, is kept after the shim.

## One network per epoch

A gang’s network lives exactly as long as one **epoch** of the gang: one attempt at running it with a fixed set of members. Each epoch gets a fresh mesh identity with single-use join keys that expire after 15 minutes, and its members can reach only each other. When Nodus restarts the gang, for example to replace a lost member, it fences the old epoch first: the old network is torn down, and a member left behind on it can no longer reach the new gang, nor can the new gang reach it. Members of other gangs, in your org or anyone else’s, are never reachable.

## The probe

Before your command starts on any member, every member checks its path to every other member. Nodus bills this time as boot time.

1. **Readiness.** Each member opens three fresh TCP connections to every peer, on ephemeral ports, within 60 seconds, through the same path your processes will use.
2. **Measurement.** For each pair, the member records the round-trip time, a 10-second TCP throughput sample and the path class it actually got: `direct`, or `derp:<region>` when the pair goes through a relay.
3. **Shim self-test** (`Relayed` only). The member runs a check program with your image’s own loader to confirm that the shim loads.

The slowest pair is published in `status.gang.probe` as `tcpGbps` and `rttMs`, with its path, and sets the `NetworkQualified` condition. A pair that cannot connect, or a path worse than your `transport` allows, fails the epoch and Nodus places the gang again; that time is not billed to you. The one exception is an image whose programs cannot load the shim: placing the gang elsewhere cannot fix it, so the Job fails at once with `ShimNotLoaded` and the time is billed (see [caveats](https://nodus-platform-site.pages.dev/docs/concepts/gang-networking/#caveats)).

`tcpGbps` is a TCP throughput sample between two members, not NCCL bus bandwidth. Read it as an upper bound on what one connection between that pair can carry.

Terminal window

```console
$ nodus get job llama-ft -o jsonpath='{.status.gang.probe}'
{"tcpGbps":0.41,"rttMs":38.2,"path":"derp:nodus-us","measuredTime":"…"}
```

## Measured throughput

Nodus publishes only measured numbers: the probe’s TCP throughput and RTT, and NCCL all-reduce bus bandwidth from a separate benchmark on 2 × 1 H100. Nodus places a gang only on combinations of offerings it has qualified for the path, and a row appears here once its qualification run passes.

|Path|Members|TCP throughput (probe)|RTT|NCCL bus bandwidth|
|-|-|-|-|-|
|`Private`|One provider region|Not yet published|Not yet published|Not yet published|
|`Direct`|VMs in different providers|Not yet published|Not yet published|Not yet published|
|`Relayed`|A container-only offering with a VM|Not yet published|Not yet published|Not yet published|

Until the benchmark publishes, treat `Relayed` as low bandwidth: the scheduler assumes it is four times slower than `Direct` when it compares placements, and the estimate warns about it.

## Billing

`Private` and `Direct` traffic costs nothing beyond the members’ own time. `Relayed` traffic crosses Nodus relays and is billed per GiB on the `mesh-relay` line of your usage, measured on the relays as the bytes each member sends. Mesh traffic never counts as container egress. Each `Relayed` gang may send up to 1 Gbit/s through the relays, split evenly across its members; the relays are shared and best effort within that limit.

## Caveats

* **Synchronous training across providers is bound by the WAN.** Expect cross-provider data-parallel training to be limited by network bandwidth and latency, not by the GPUs. `Relayed` suits small models, algorithms that communicate little, reinforcement learning with separate rollouts, and getting N GPUs now wherever they are.
* **`Relayed` needs dynamically linked glibc programs.** The shim is a preloaded library, so it loads only into programs that use the system’s glibc loader. Launchers built on musl (Alpine images) or linked statically never load it and cannot reach their peers. The probe detects this and the Job fails with `ShimNotLoaded`. Fix it by using a glibc-based image (any Debian, Ubuntu or CUDA image works), or by setting `transport: Direct`.
* **`Relayed` is low bandwidth.** All gang traffic crosses the relays, under the per-gang limit above. Nodus makes no throughput claim for `Relayed` until the benchmark is published.
* **Advertised addresses must be distinct and exclusive.** On `Relayed`, each member is reached at its own `eth0` address, so no two members of a gang may share one, and no two running `Relayed` gangs may route the same address. When placement produces such a collision, Nodus places the members again; it shows as `IPCollision` in the `GangReady` condition while it does.
* **Only connections to peers go through the mesh.** On `Relayed`, connections to other members are redirected; UDP and connections to any other address are not. Tools that need UDP between members, or connect to peers through a hostname that does not resolve to their advertised address, do not work on `Relayed`.
* **The probe measures TCP.** A good `tcpGbps` does not guarantee NCCL performance; NCCL’s numbers come only from the benchmark in the table above.
