A private LLM platform on 32 GPUs
How we would size a GPU platform for a financial-services group that wants an internal AI assistant for 2,000 staff, plus regular fine-tuning, without sending data outside its own infrastructure.
This is a composite example built from requirements common in regulated enterprises. It is not a specific client engagement, and the figures are design estimates for the scenario, not measured results.
Useful AI, but no data leaving the building
The group wanted an assistant that could answer questions on policies, products and internal procedures, and summarise long documents. Its risk team ruled out sending customer or internal data to an external AI service. It also wanted to fine-tune models on its own terminology every quarter.
Working out how many GPUs
| Workload | Assumption | GPUs |
|---|---|---|
| Assistant (70B-class model, 8-bit) | Peak 200 people asking at once, up to 8,000 tokens of context each | 8 GPUs: 4 replicas of 2 x H200 |
| Quick tasks (8B-class model) | Classification, extraction, short answers | 4 GPUs, shared |
| Fine-tuning | LoRA on the 70B model, full fine-tune of the 8B model, quarterly | 8 GPUs (one server), idle between runs for burst inference |
| Headroom and failover | One server can fail without losing the service | 12 GPUs |
| Total | 32 GPUs in 4 servers |
KV cache for 200 users at 8,000 tokens on a 70B-class model at 8-bit is roughly 250 GB, spread across the replicas. H200's 141 GB per GPU leaves room for it alongside the weights.
What gets built
| Component | Design |
|---|---|
| GPU servers | 4 x 8-GPU H200 servers |
| GPU fabric | One 64-port 400G InfiniBand switch: 32 ports used, room to double |
| Front-end and storage network | 100/200 Gb/s Ethernet, separate from the GPU fabric |
| Storage | About 300 TB all-flash parallel file system, plus object storage for documents and models |
| Platform | Kubernetes with GPU Operator, vLLM model servers, API gateway with per-department quotas |
| Facility | ~47 kW across 2 racks at ~24 kW each, with rear-door heat exchangers |
What we would flag
- H200 costs more per GPU than H100. Its extra memory means fewer GPUs per model replica, so the total can come out similar. Compare complete configurations, not unit prices.
- The data centre decides the timeline. If the facility needs new power or chilled water for two 24 kW racks, that work starts before the order is placed.
- Peak usage drives cost. Most demand comes in the first hours of the working day. A short queue at peak can save several GPUs.
- The model is only half of it. Answers about internal policies need good retrieval from internal documents. See RAG solutions.
Planning a GPU cluster?
Tell us the models you want to run or train, how many users, and where it will be hosted. We will come back with a first sizing: GPUs, servers, network, storage, and the power and cooling your data centre will need.