Skip to content

Gang networking

View Markdown

Beta: this feature may change before general availability

A distributed Job runs as a gang: one member per machine, every member started together, sometimes on machines from different providers. Before your command starts, Nodus joins the members into a private network that belongs to that gang alone. Each rank gets the addresses of the others in its environment, so torchrun, Ray and plain torch.distributed work without any networking code of yours.

You choose two things in spec.distributed:

  • network: how close the members must be. Colocated keeps them in one provider region, Regional allows any provider inside one region class, and Global allows anywhere.
  • transport: which paths you accept. Direct (the default) accepts only Private and Direct paths. Auto also accepts Relayed paths, which lets members on machines without a kernel network device join.

Nodus picks the path. Every member of a gang uses the same path class, so every rank sees the same addressing and the same NCCL settings.

Path When Nodus uses it How traffic flows Encryption Bandwidth
Private Every member is in one provider region with a private network The provider’s private network, opened only between the gang’s members None added (the provider’s network) Provider native
Direct Every member can create a WireGuard device and the probe finds a direct path between every pair WireGuard between the members, on a network device named nodus0 WireGuard Measured per pair (see measured throughput)
Relayed A member has no network device of its own (container-only offerings) and transport: Auto WireGuard in user space, carried through Nodus relays WireGuard end to end; relays see only ciphertext Low: small models only

Members in the same provider region talk over the provider’s private network. Nodus opens the rendezvous and NCCL ports between the gang’s members only and closes them when the gang ends. The advertised addresses are the members’ private IPs.

Each member runs a WireGuard endpoint next to your container, never inside it: your container gets no extra privileges. The members find each other through the Nodus mesh control and connect peer to peer, punching through NAT where needed. The advertised addresses are the members’ mesh addresses (100.64.0.0/10), and NCCL and Gloo use the nodus0 interface.

Direct means direct. If the probe finds that one pair can only connect through a relay, a transport: Direct gang is placed again without that pair, and a transport: Auto gang continues as Relayed. If a running pair loses its direct path, its traffic falls back to a relay without interrupting training, and the Job’s NetworkDegraded condition becomes true.

Container-only offerings give your container no network device, no CAP_NET_ADMIN and no UDP. Their members can still join a gang: each runs WireGuard in user space and sends peer traffic through two relays that Nodus operates. To reach that WireGuard endpoint, your processes start with a small preloaded library, libnodus-netshim.so, which redirects only connections to the other members of your gang. Everything else, such as downloads, datasets, object storage and model APIs, goes out directly as usual.

The advertised addresses are the members’ own eth0 addresses, so MASTER_ADDR, the torchrun rendezvous and the listeners NCCL opens on ephemeral ports all work unchanged, in both directions. The relays mesh with each other: if one goes down, members move to the other within seconds and the TCP connections inside WireGuard survive.

The path decides a few variables; the rest of the distributed environment is the same on every path.

Variable Private Direct Relayed
NODUS_GANG_TRANSPORT private direct relayed
MASTER_ADDR, PET_RDZV_ENDPOINT, NODUS_NODE_IPS Private IPs Mesh addresses eth0 addresses
NCCL_SOCKET_IFNAME, GLOO_SOCKET_IFNAME The private interface nodus0 eth0
LD_PRELOAD, NODUS_NETSHIM_PEERS, NODUS_NETSHIM_SOCKS Not set Not set The shim first, then your image’s own LD_PRELOAD

Your own NCCL_* values, such as NCCL_DEBUG=INFO, win, except NCCL_NET, NCCL_SOCKET_IFNAME, NCCL_SOCKET_FAMILY and NCCL_IB_DISABLE, which the path fixes. A spec LD_PRELOAD is rejected when transport: Auto, because Nodus needs to place the shim first on a Relayed path. An LD_PRELOAD set in the image, such as jemalloc or tcmalloc, is kept after the shim.

A gang’s network lives exactly as long as one epoch of the gang: one attempt at running it with a fixed set of members. Each epoch gets a fresh mesh identity with single-use join keys that expire after 15 minutes, and its members can reach only each other. When Nodus restarts the gang, for example to replace a lost member, it fences the old epoch first: the old network is torn down, and a member left behind on it can no longer reach the new gang, nor can the new gang reach it. Members of other gangs, in your org or anyone else’s, are never reachable.

Before your command starts on any member, every member checks its path to every other member. Nodus bills this time as boot time.

  1. Readiness. Each member opens three fresh TCP connections to every peer, on ephemeral ports, within 60 seconds, through the same path your processes will use.
  2. Measurement. For each pair, the member records the round-trip time, a 10-second TCP throughput sample and the path class it actually got: direct, or derp:<region> when the pair goes through a relay.
  3. Shim self-test (Relayed only). The member runs a check program with your image’s own loader to confirm that the shim loads.

The slowest pair is published in status.gang.probe as tcpGbps and rttMs, with its path, and sets the NetworkQualified condition. A pair that cannot connect, or a path worse than your transport allows, fails the epoch and Nodus places the gang again; that time is not billed to you. The one exception is an image whose programs cannot load the shim: placing the gang elsewhere cannot fix it, so the Job fails at once with ShimNotLoaded and the time is billed (see caveats).

tcpGbps is a TCP throughput sample between two members, not NCCL bus bandwidth. Read it as an upper bound on what one connection between that pair can carry.

Terminal window
$ nodus get job llama-ft -o jsonpath='{.status.gang.probe}'
{"tcpGbps":0.41,"rttMs":38.2,"path":"derp:nodus-us","measuredTime":"…"}

Nodus publishes only measured numbers: the probe’s TCP throughput and RTT, and NCCL all-reduce bus bandwidth from a separate benchmark on 2 × 1 H100. Nodus places a gang only on combinations of offerings it has qualified for the path, and a row appears here once its qualification run passes.

Path Members TCP throughput (probe) RTT NCCL bus bandwidth
Private One provider region Not yet published Not yet published Not yet published
Direct VMs in different providers Not yet published Not yet published Not yet published
Relayed A container-only offering with a VM Not yet published Not yet published Not yet published

Until the benchmark publishes, treat Relayed as low bandwidth: the scheduler assumes it is four times slower than Direct when it compares placements, and the estimate warns about it.

Private and Direct traffic costs nothing beyond the members’ own time. Relayed traffic crosses Nodus relays and is billed per GiB on the mesh-relay line of your usage, measured on the relays as the bytes each member sends. Mesh traffic never counts as container egress. Each Relayed gang may send up to 1 Gbit/s through the relays, split evenly across its members; the relays are shared and best effort within that limit.

  • Synchronous training across providers is bound by the WAN. Expect cross-provider data-parallel training to be limited by network bandwidth and latency, not by the GPUs. Relayed suits small models, algorithms that communicate little, reinforcement learning with separate rollouts, and getting N GPUs now wherever they are.
  • Relayed needs dynamically linked glibc programs. The shim is a preloaded library, so it loads only into programs that use the system’s glibc loader. Launchers built on musl (Alpine images) or linked statically never load it and cannot reach their peers. The probe detects this and the Job fails with ShimNotLoaded. Fix it by using a glibc-based image (any Debian, Ubuntu or CUDA image works), or by setting transport: Direct.
  • Relayed is low bandwidth. All gang traffic crosses the relays, under the per-gang limit above. Nodus makes no throughput claim for Relayed until the benchmark is published.
  • Advertised addresses must be distinct and exclusive. On Relayed, each member is reached at its own eth0 address, so no two members of a gang may share one, and no two running Relayed gangs may route the same address. When placement produces such a collision, Nodus places the members again; it shows as IPCollision in the GangReady condition while it does.
  • Only connections to peers go through the mesh. On Relayed, connections to other members are redirected; UDP and connections to any other address are not. Tools that need UDP between members, or connect to peers through a hostname that does not resolve to their advertised address, do not work on Relayed.
  • The probe measures TCP. A good tcpGbps does not guarantee NCCL performance; NCCL’s numbers come only from the benchmark in the table above.