GPU network fabric
When a training job spans several servers, the GPUs exchange results after every step. The network that carries this traffic decides whether 32 GPUs behave like 32, or like 12. It is a separate network, built only for GPU-to-GPU traffic.
Rail-optimised design
- Each GPU has its own network port, so an 8-GPU server has 8 fabric ports: 3.2 Tb/s per server.
- Leaf switches are grouped into rails. GPU 1 of every server goes to rail 1, GPU 2 to rail 2, and so on.
- Every leaf connects to every spine, so any GPU can still reach any other.
- The fabric is non-blocking: as much bandwidth up to the spines as down to the GPUs.
InfiniBand or Ethernet?
| InfiniBand (NDR 400G) | RoCE Ethernet (400G) | |
|---|---|---|
| How it works | Purpose-built for HPC, lossless by design | Standard Ethernet tuned to be lossless (PFC, ECN) |
| Strengths | Proven at large scale for training, predictable performance, simpler to tune | Familiar to network teams, wider vendor choice, one technology across the data centre |
| Watch out for | Separate skills and tooling, mostly one vendor | Needs careful configuration and testing to match InfiniBand under load |
| Usually fits | Dedicated training clusters | Mixed environments and teams with strong Ethernet skills |
Both work. The right choice depends more on your team and your existing network than on benchmark numbers. For inference-only platforms, a high-speed fabric is often not needed at all.
How many switches?
| Cluster size | GPU ports | Typical fabric |
|---|---|---|
| 4 servers (32 GPUs) | 32 x 400G | A single 64-port 400G switch can connect them all, with room to grow |
| 16 servers (128 GPUs) | 128 x 400G | Rail-optimised leaves plus spines in two tiers |
| 32 servers (256 GPUs) | 256 x 400G | A standard "scalable unit": 8 rail leaves and a spine layer |
Based on 64-port 400G switches such as the NVIDIA Quantum-2 family. Cable count and lengths matter as much as switch count: plan the rack layout at the same time.
The other networks
- Storage network: 100 to 400 Gb/s Ethernet or InfiniBand to the parallel file system, separate from the GPU fabric.
- In-band management: ordinary Ethernet for logins, scheduling and software updates.
- Out-of-band management: a separate network to every server's BMC, so failed servers can be reached and restarted.
More AI infrastructure guides: Training cluster · Inference platform · Storage for AI · Power and cooling · FAQ · Use case: private LLM platform
Planning a GPU cluster?
Tell us the models you want to run or train, how many users, and where it will be hosted. We will come back with a first sizing: GPUs, servers, network, storage, and the power and cooling your data centre will need.