GPU and AI infrastructure questions, answered
Short answers to the questions we hear most often when organisations plan GPU infrastructure.
Should we buy GPUs or rent them from a GPU cloud?
Rent while you learn your real usage, or for short projects. Buy when you have steady, long-term demand, or when data must stay on your premises. Many organisations do both: own the steady base, rent for peaks.
H100, H200 or L40S?
H100 and H200 are built for training and large-model serving; H200 has much more memory (141 GB vs 80 GB), which helps large models and long contexts. L40S is a cheaper PCIe GPU that serves small and mid-size models well but is not designed for multi-server training.
Do we need InfiniBand?
For training across several servers, you need a high-speed GPU fabric: InfiniBand or well-tuned RoCE Ethernet. For inference where each model fits in one server, standard data-centre networking is usually enough.
How long do GPU servers take to arrive?
It varies with demand and model. Lead times of several weeks to a few months are common. Facility upgrades for power and cooling can take longer, so start that check first.
Can our existing data centre host a GPU cluster?
Often only partly. One 8-GPU server draws about 10 kW, more than many enterprise racks are built for. We check power, cooling and floor loading before anything is ordered. See power and cooling.
How do we keep GPUs busy?
A scheduler that understands GPUs, quotas per team, GPU sharing (MIG) for small workloads, and monitoring that shows utilisation per team. Clusters without these often run far below capacity.
Can we run training and inference on the same cluster?
Yes, with care. Inference needs steady capacity during working hours; training can use what is left, especially at night. Kubernetes or a scheduler with priorities and pre-emption makes this work.
What do you need from us to start?
The models you plan to use, training or inference, rough user numbers, and where the hardware will live. We will start from that and the data-centre details.
More AI infrastructure guides: Training cluster · Inference platform · GPU network fabric · Storage for AI · Power and cooling · FAQ · Use case: private LLM platform
Planning a GPU cluster?
Tell us the models you want to run or train, how many users, and where it will be hosted. We will come back with a first sizing: GPUs, servers, network, storage, and the power and cooling your data centre will need.