Every GPU of a machine for torchrun
- A Job with several GPUs on one machine (
--gpu A6000:2,--gpu H100:8) setsPET_NPROC_PER_NODEto the GPU count, sotorchrun train.pystarts one process per GPU. A--nproc-per-nodeyou pass yourself wins. See Use every GPU of one machine. /dev/shmin a Job’s container is half its memory limit, or half the machine’s memory, instead of 64 MiB, so NCCL between GPUs and PyTorch data-loader workers no longer run out of shared memory.