Gang networking
View MarkdownBeta: this feature may change before general availability
A distributed Job runs as a gang: one member per machine, every member started together, sometimes on machines
from different providers. Before your command starts, Nodus joins the members into a private network that belongs
to that gang alone. Each rank gets the addresses of the others in its environment, so torchrun, Ray and plain
torch.distributed work without any networking code of yours.
You choose two things in spec.distributed:
network: how close the members must be.Colocatedkeeps them in one provider region,Regionalallows any provider inside one region class, andGlobalallows anywhere.transport: which paths you accept.Direct(the default) accepts onlyPrivateandDirectpaths.Autoalso acceptsRelayedpaths, which lets members on machines without a kernel network device join.
Nodus picks the path. Every member of a gang uses the same path class, so every rank sees the same addressing and the same NCCL settings.
Path classes
Section titled “Path classes”| Path | When Nodus uses it | How traffic flows | Encryption | Bandwidth |
|---|---|---|---|---|
Private |
Every member is in one provider region with a private network | The provider’s private network, opened only between the gang’s members | None added (the provider’s network) | Provider native |
Direct |
Every member can create a WireGuard device and the probe finds a direct path between every pair | WireGuard between the members, on a network device named nodus0 |
WireGuard | Measured per pair (see measured throughput) |
Relayed |
A member has no network device of its own (container-only offerings) and transport: Auto |
WireGuard in user space, carried through Nodus relays | WireGuard end to end; relays see only ciphertext | Low: small models only |
Private
Section titled “Private”Members in the same provider region talk over the provider’s private network. Nodus opens the rendezvous and NCCL ports between the gang’s members only and closes them when the gang ends. The advertised addresses are the members’ private IPs.
Direct
Section titled “Direct”Each member runs a WireGuard endpoint next to your container, never inside it: your container gets no extra
privileges. The members find each other through the Nodus mesh control and connect peer to peer, punching through
NAT where needed. The advertised addresses are the members’ mesh addresses (100.64.0.0/10), and NCCL and Gloo
use the nodus0 interface.
Direct means direct. If the probe finds that one pair can only connect through a relay, a transport: Direct
gang is placed again without that pair, and a transport: Auto gang continues as Relayed. If a running pair
loses its direct path, its traffic falls back to a relay without interrupting training, and the Job’s
NetworkDegraded condition becomes true.
Relayed
Section titled “Relayed”Container-only offerings give your container no network device, no CAP_NET_ADMIN and no UDP. Their
members can still join a gang: each runs WireGuard in user space and sends peer traffic through two relays that
Nodus operates. To reach that WireGuard endpoint, your processes start with a small preloaded library,
libnodus-netshim.so, which redirects only connections to the other members of your gang. Everything else,
such as downloads, datasets, object storage and model APIs, goes out directly as usual.
The advertised addresses are the members’ own eth0 addresses, so MASTER_ADDR, the torchrun rendezvous and
the listeners NCCL opens on ephemeral ports all work unchanged, in both directions. The relays mesh with each
other: if one goes down, members move to the other within seconds and the TCP connections inside WireGuard
survive.
What every rank sees
Section titled “What every rank sees”The path decides a few variables; the rest of the distributed environment is the same on every path.
| Variable | Private |
Direct |
Relayed |
|---|---|---|---|
NODUS_GANG_TRANSPORT |
private |
direct |
relayed |
MASTER_ADDR, PET_RDZV_ENDPOINT, NODUS_NODE_IPS |
Private IPs | Mesh addresses | eth0 addresses |
NCCL_SOCKET_IFNAME, GLOO_SOCKET_IFNAME |
The private interface | nodus0 |
eth0 |
LD_PRELOAD, NODUS_NETSHIM_PEERS, NODUS_NETSHIM_SOCKS |
Not set | Not set | The shim first, then your image’s own LD_PRELOAD |
Your own NCCL_* values, such as NCCL_DEBUG=INFO, win, except NCCL_NET, NCCL_SOCKET_IFNAME,
NCCL_SOCKET_FAMILY and NCCL_IB_DISABLE, which the path fixes. A spec LD_PRELOAD is rejected when
transport: Auto, because Nodus needs to place the shim first on a Relayed path. An LD_PRELOAD set in the
image, such as jemalloc or tcmalloc, is kept after the shim.
One network per epoch
Section titled “One network per epoch”A gang’s network lives exactly as long as one epoch of the gang: one attempt at running it with a fixed set of members. Each epoch gets a fresh mesh identity with single-use join keys that expire after 15 minutes, and its members can reach only each other. When Nodus restarts the gang, for example to replace a lost member, it fences the old epoch first: the old network is torn down, and a member left behind on it can no longer reach the new gang, nor can the new gang reach it. Members of other gangs, in your org or anyone else’s, are never reachable.
The probe
Section titled “The probe”Before your command starts on any member, every member checks its path to every other member. Nodus bills this time as boot time.
- Readiness. Each member opens three fresh TCP connections to every peer, on ephemeral ports, within 60 seconds, through the same path your processes will use.
- Measurement. For each pair, the member records the round-trip time, a 10-second TCP throughput sample and
the path class it actually got:
direct, orderp:<region>when the pair goes through a relay. - Shim self-test (
Relayedonly). The member runs a check program with your image’s own loader to confirm that the shim loads.
The slowest pair is published in status.gang.probe as tcpGbps and rttMs, with its path, and sets the
NetworkQualified condition. A pair that cannot connect, or a path worse than your transport allows, fails the
epoch and Nodus places the gang again; that time is not billed to you. The one exception is an image whose
programs cannot load the shim: placing the gang elsewhere cannot fix it, so the Job fails at once with
ShimNotLoaded and the time is billed (see caveats).
tcpGbps is a TCP throughput sample between two members, not NCCL bus bandwidth. Read it as an upper bound on
what one connection between that pair can carry.
$ nodus get job llama-ft -o jsonpath='{.status.gang.probe}'{"tcpGbps":0.41,"rttMs":38.2,"path":"derp:nodus-us","measuredTime":"…"}Measured throughput
Section titled “Measured throughput”Nodus publishes only measured numbers: the probe’s TCP throughput and RTT, and NCCL all-reduce bus bandwidth from a separate benchmark on 2 × 1 H100. Nodus places a gang only on combinations of offerings it has qualified for the path, and a row appears here once its qualification run passes.
| Path | Members | TCP throughput (probe) | RTT | NCCL bus bandwidth |
|---|---|---|---|---|
Private |
One provider region | Not yet published | Not yet published | Not yet published |
Direct |
VMs in different providers | Not yet published | Not yet published | Not yet published |
Relayed |
A container-only offering with a VM | Not yet published | Not yet published | Not yet published |
Until the benchmark publishes, treat Relayed as low bandwidth: the scheduler assumes it is four times slower
than Direct when it compares placements, and the estimate warns about it.
Billing
Section titled “Billing”Private and Direct traffic costs nothing beyond the members’ own time. Relayed traffic crosses Nodus relays
and is billed per GiB on the mesh-relay line of your usage, measured on the relays as the bytes each member
sends. Mesh traffic never counts as container egress. Each Relayed gang may send up to 1 Gbit/s through the
relays, split evenly across its members; the relays are shared and best effort within that limit.
Caveats
Section titled “Caveats”- Synchronous training across providers is bound by the WAN. Expect cross-provider data-parallel training to
be limited by network bandwidth and latency, not by the GPUs.
Relayedsuits small models, algorithms that communicate little, reinforcement learning with separate rollouts, and getting N GPUs now wherever they are. Relayedneeds dynamically linked glibc programs. The shim is a preloaded library, so it loads only into programs that use the system’s glibc loader. Launchers built on musl (Alpine images) or linked statically never load it and cannot reach their peers. The probe detects this and the Job fails withShimNotLoaded. Fix it by using a glibc-based image (any Debian, Ubuntu or CUDA image works), or by settingtransport: Direct.Relayedis low bandwidth. All gang traffic crosses the relays, under the per-gang limit above. Nodus makes no throughput claim forRelayeduntil the benchmark is published.- Advertised addresses must be distinct and exclusive. On
Relayed, each member is reached at its owneth0address, so no two members of a gang may share one, and no two runningRelayedgangs may route the same address. When placement produces such a collision, Nodus places the members again; it shows asIPCollisionin theGangReadycondition while it does. - Only connections to peers go through the mesh. On
Relayed, connections to other members are redirected; UDP and connections to any other address are not. Tools that need UDP between members, or connect to peers through a hostname that does not resolve to their advertised address, do not work onRelayed. - The probe measures TCP. A good
tcpGbpsdoes not guarantee NCCL performance; NCCL’s numbers come only from the benchmark in the table above.