Home / AI infrastructure / GPU network fabric
AI infrastructure guide

GPU network fabric

When a training job spans several servers, the GPUs exchange results after every step. The network that carries this traffic decides whether 32 GPUs behave like 32, or like 12. It is a separate network, built only for GPU-to-GPU traffic.

Speed per GPU400 Gb/s
OptionsInfiniBand or RoCE Ethernet
LayoutRail-optimised
Needed forMulti-server training
How it is wired

Rail-optimised design

Scroll sideways to see the whole diagram →
Spine 1Spine 2Spine 3Spine 4Rail 1 leafRail 2 leafRail 3 leafRail 4 leafRail 5 leafRail 6 leafRail 7 leafRail 8 leaf8-GPU server 1NVLink inside the server8-GPU server 2NVLink inside the server8-GPU server 3NVLink inside the server8-GPU server 4NVLink inside the serverGPU 1 of every server connects to rail 1, GPU 2 to rail 2, and so on. Every leaf connects to every spine.
Each colour is a rail. Traffic between the same-numbered GPUs on different servers crosses just one leaf switch, which keeps latency low and predictable.
  • Each GPU has its own network port, so an 8-GPU server has 8 fabric ports: 3.2 Tb/s per server.
  • Leaf switches are grouped into rails. GPU 1 of every server goes to rail 1, GPU 2 to rail 2, and so on.
  • Every leaf connects to every spine, so any GPU can still reach any other.
  • The fabric is non-blocking: as much bandwidth up to the spines as down to the GPUs.

InfiniBand or Ethernet?

InfiniBand (NDR 400G)RoCE Ethernet (400G)
How it worksPurpose-built for HPC, lossless by designStandard Ethernet tuned to be lossless (PFC, ECN)
StrengthsProven at large scale for training, predictable performance, simpler to tuneFamiliar to network teams, wider vendor choice, one technology across the data centre
Watch out forSeparate skills and tooling, mostly one vendorNeeds careful configuration and testing to match InfiniBand under load
Usually fitsDedicated training clustersMixed environments and teams with strong Ethernet skills

Both work. The right choice depends more on your team and your existing network than on benchmark numbers. For inference-only platforms, a high-speed fabric is often not needed at all.

How many switches?

Cluster sizeGPU portsTypical fabric
4 servers (32 GPUs)32 x 400GA single 64-port 400G switch can connect them all, with room to grow
16 servers (128 GPUs)128 x 400GRail-optimised leaves plus spines in two tiers
32 servers (256 GPUs)256 x 400GA standard "scalable unit": 8 rail leaves and a spine layer

Based on 64-port 400G switches such as the NVIDIA Quantum-2 family. Cable count and lengths matter as much as switch count: plan the rack layout at the same time.

The other networks

  • Storage network: 100 to 400 Gb/s Ethernet or InfiniBand to the parallel file system, separate from the GPU fabric.
  • In-band management: ordinary Ethernet for logins, scheduling and software updates.
  • Out-of-band management: a separate network to every server's BMC, so failed servers can be reached and restarted.

Planning a GPU cluster?

Tell us the models you want to run or train, how many users, and where it will be hosted. We will come back with a first sizing: GPUs, servers, network, storage, and the power and cooling your data centre will need.