Home / AI infrastructure / Use case: private LLM platform
Illustrative scenario

A private LLM platform on 32 GPUs

How we would size a GPU platform for a financial-services group that wants an internal AI assistant for 2,000 staff, plus regular fine-tuning, without sending data outside its own infrastructure.

This is a composite example built from requirements common in regulated enterprises. It is not a specific client engagement, and the figures are design estimates for the scenario, not measured results.

Users~2,000 staff
Models70B and 8B class
GPUs32 x H200
Power~47 kW, 2 racks
The situation

Useful AI, but no data leaving the building

The group wanted an assistant that could answer questions on policies, products and internal procedures, and summarise long documents. Its risk team ruled out sending customer or internal data to an external AI service. It also wanted to fine-tune models on its own terminology every quarter.

Sizing

Working out how many GPUs

WorkloadAssumptionGPUs
Assistant (70B-class model, 8-bit)Peak 200 people asking at once, up to 8,000 tokens of context each8 GPUs: 4 replicas of 2 x H200
Quick tasks (8B-class model)Classification, extraction, short answers4 GPUs, shared
Fine-tuningLoRA on the 70B model, full fine-tune of the 8B model, quarterly8 GPUs (one server), idle between runs for burst inference
Headroom and failoverOne server can fail without losing the service12 GPUs
Total32 GPUs in 4 servers

KV cache for 200 users at 8,000 tokens on a 70B-class model at 8-bit is roughly 250 GB, spread across the replicas. H200's 141 GB per GPU leaves room for it alongside the weights.

The design

What gets built

ComponentDesign
GPU servers4 x 8-GPU H200 servers
GPU fabricOne 64-port 400G InfiniBand switch: 32 ports used, room to double
Front-end and storage network100/200 Gb/s Ethernet, separate from the GPU fabric
StorageAbout 300 TB all-flash parallel file system, plus object storage for documents and models
PlatformKubernetes with GPU Operator, vLLM model servers, API gateway with per-department quotas
Facility~47 kW across 2 racks at ~24 kW each, with rear-door heat exchangers
Trade-offs

What we would flag

  • H200 costs more per GPU than H100. Its extra memory means fewer GPUs per model replica, so the total can come out similar. Compare complete configurations, not unit prices.
  • The data centre decides the timeline. If the facility needs new power or chilled water for two 24 kW racks, that work starts before the order is placed.
  • Peak usage drives cost. Most demand comes in the first hours of the working day. A short queue at peak can save several GPUs.
  • The model is only half of it. Answers about internal policies need good retrieval from internal documents. See RAG solutions.

Planning a GPU cluster?

Tell us the models you want to run or train, how many users, and where it will be hosted. We will come back with a first sizing: GPUs, servers, network, storage, and the power and cooling your data centre will need.