Home / AI infrastructure / Training cluster
AI infrastructure guide

GPU training cluster

Training a model across several GPU servers turns them into one machine: every GPU works on its own slice of data, then all of them share results before the next step. That only works if the servers, the fabric between them and the storage behind them are designed together.

Typical building block8-GPU server
GPU fabric400 Gb/s per GPU
SchedulerSlurm or Kubernetes
Must haveRegular checkpoints
See it working

A training job, and what happens when a server fails

Watch the cycle: load data (storage lights up), compute (GPUs light up), synchronise (fabric lights up). Then fail a node. Because every GPU depends on every other GPU at each step, the whole job stops and restarts from the last checkpoint. Try different checkpoint intervals and compare how much work is lost.

Reference architecture

How the pieces fit together

Scroll sideways to see the whole diagram →
GPU computeScheduling, storage and monitoringData scientists and pipelinessubmit training jobsGPU servers 1-28x H100 or H200 eachGPU servers 3-48x H100 or H200 eachNVLink and NVSwitchGPU-to-GPU inside each serverCompute fabric400 Gb/s per GPU, InfiniBand or RoCEStorage and management networkskept separate from the compute fabricSchedulerSlurm, or Kubernetes with GPU OperatorParallel file systemdatasets and checkpointsObject storageraw data and model registryMonitoringDCGM, Prometheus, GrafanaPower and cooling~10 kW per 8-GPU server123
Numbered lines show the main flows. Compute, storage and management traffic each have their own network.
PartWhat it does
1 SchedulingThe scheduler places each job on whole servers, so GPUs that need to talk are close together on the fabric.
Inside a serverNVLink and NVSwitch connect the 8 GPUs inside a server at very high speed. Most communication stays here when a job fits in one server.
Between serversThe compute fabric gives each GPU its own 400 Gb/s port. This is what lets 32 or 256 GPUs train as one.
2 Data and checkpointsTraining data is read from, and checkpoints are written to, a parallel file system over a separate storage network.
3 StagingRaw data and finished models live in cheaper object storage and are staged to the fast tier when needed.
MonitoringGPU temperature, errors and utilisation are tracked per GPU, because one failing GPU slows or stops the whole job.

Getting it right

  • Size checkpoints, not just storage capacity. A checkpoint of a 70-billion-parameter model with optimizer state is around 1 TB. If writing it takes 10 minutes, GPUs are idle for those 10 minutes every time.
  • Plan for failure. In large clusters, a GPU, cable or server will fail during long runs. Checkpoint often enough that a failure costs minutes, not days, and keep a spare server ready.
  • Burn in before handover. Run stress tests and fabric bandwidth tests (such as NCCL tests) on every server. Weak links found later cost much more.
  • Keep the fabric only for GPUs. Storage and management traffic on the compute fabric slows training in ways that are hard to diagnose.
  • Choose the scheduler for how people work. Research teams used to batch jobs often prefer Slurm. Platform teams running mixed workloads lean towards Kubernetes. Some clusters run both.

Typical tools

LayerCommon choices
ServersNVIDIA HGX H100, H200 or B200-based servers from major OEMs; DGX systems
SchedulingSlurm, Kubernetes with NVIDIA GPU Operator, Run:ai-style quota management
CommunicationNCCL over InfiniBand or RoCE
FrameworksPyTorch with FSDP or DeepSpeed, NVIDIA NeMo
MonitoringNVIDIA DCGM exporter, Prometheus, Grafana

Planning a GPU cluster?

Tell us the models you want to run or train, how many users, and where it will be hosted. We will come back with a first sizing: GPUs, servers, network, storage, and the power and cooling your data centre will need.