GPU training cluster
Training a model across several GPU servers turns them into one machine: every GPU works on its own slice of data, then all of them share results before the next step. That only works if the servers, the fabric between them and the storage behind them are designed together.
A training job, and what happens when a server fails
Watch the cycle: load data (storage lights up), compute (GPUs light up), synchronise (fabric lights up). Then fail a node. Because every GPU depends on every other GPU at each step, the whole job stops and restarts from the last checkpoint. Try different checkpoint intervals and compare how much work is lost.
How the pieces fit together
| Part | What it does |
|---|---|
| 1 Scheduling | The scheduler places each job on whole servers, so GPUs that need to talk are close together on the fabric. |
| Inside a server | NVLink and NVSwitch connect the 8 GPUs inside a server at very high speed. Most communication stays here when a job fits in one server. |
| Between servers | The compute fabric gives each GPU its own 400 Gb/s port. This is what lets 32 or 256 GPUs train as one. |
| 2 Data and checkpoints | Training data is read from, and checkpoints are written to, a parallel file system over a separate storage network. |
| 3 Staging | Raw data and finished models live in cheaper object storage and are staged to the fast tier when needed. |
| Monitoring | GPU temperature, errors and utilisation are tracked per GPU, because one failing GPU slows or stops the whole job. |
Getting it right
- Size checkpoints, not just storage capacity. A checkpoint of a 70-billion-parameter model with optimizer state is around 1 TB. If writing it takes 10 minutes, GPUs are idle for those 10 minutes every time.
- Plan for failure. In large clusters, a GPU, cable or server will fail during long runs. Checkpoint often enough that a failure costs minutes, not days, and keep a spare server ready.
- Burn in before handover. Run stress tests and fabric bandwidth tests (such as NCCL tests) on every server. Weak links found later cost much more.
- Keep the fabric only for GPUs. Storage and management traffic on the compute fabric slows training in ways that are hard to diagnose.
- Choose the scheduler for how people work. Research teams used to batch jobs often prefer Slurm. Platform teams running mixed workloads lean towards Kubernetes. Some clusters run both.
Typical tools
| Layer | Common choices |
|---|---|
| Servers | NVIDIA HGX H100, H200 or B200-based servers from major OEMs; DGX systems |
| Scheduling | Slurm, Kubernetes with NVIDIA GPU Operator, Run:ai-style quota management |
| Communication | NCCL over InfiniBand or RoCE |
| Frameworks | PyTorch with FSDP or DeepSpeed, NVIDIA NeMo |
| Monitoring | NVIDIA DCGM exporter, Prometheus, Grafana |
More AI infrastructure guides: Inference platform · GPU network fabric · Storage for AI · Power and cooling · FAQ · Use case: private LLM platform
Planning a GPU cluster?
Tell us the models you want to run or train, how many users, and where it will be hosted. We will come back with a first sizing: GPUs, servers, network, storage, and the power and cooling your data centre will need.