GPU inference platform
Inference is where a model meets its users, and where most of the GPU budget ends up over time. The goals are different from training: steady low latency, many users at once, and as many requests per GPU as possible without making anyone wait.
How requests are served
| Part | What it does |
|---|---|
| API gateway | Every request is signed in, rate-limited and counted per team, so cost can be charged back and one team cannot starve another. |
| 1 Routing | Simple questions go to a small, cheap model. Hard ones go to a large model. This alone can cut GPU needs sharply. |
| 2 Model registry | Only approved, versioned models are served. Rolling back a bad model is a configuration change, not an emergency. |
| 3 Autoscaling | Replicas are added when request queues grow and removed when they shrink, within the GPUs available. |
| 4 Monitoring | Time to first token, tokens per second, GPU utilisation and cost per team, on one dashboard. |
Choosing GPUs for inference
| Model size | Fits on | Notes |
|---|---|---|
| ~8B | One L40S (48 GB) or a slice of an H100 | Cheapest to serve. Often good enough for classification, extraction and simple chat. |
| ~70B | 2 to 4 x H100 80 GB at 16-bit; 1 to 2 x H100 at 8-bit; 1 x H200 at 8-bit | The sweet spot for many enterprise assistants. 8-bit precision usually costs little quality. |
| ~400B | One full 8-GPU H200 server at 8-bit, or more | Rarely needed for enterprise use. Consider whether a 70B model with good retrieval does the job. |
Memory is only the start. Many concurrent users with long contexts need extra memory for the KV cache. Use the GPU calculator to see the effect.
Getting it right
- Batch requests. Model servers such as vLLM serve many users on one GPU by batching continuously. Serving one request at a time wastes most of the GPU.
- Share GPUs for small models. MIG splits an H100 into up to seven isolated slices. Small models do not need a whole GPU each.
- Quantise with care. 8-bit weights roughly halve memory with little quality loss for most models. Test 4-bit on your own tasks before relying on it.
- Set latency targets per use case. A chat assistant needs a fast first token. An overnight document job does not. Size for the strictest case only where it applies.
- Plan capacity for peaks. Usage often doubles at the start of the working day. Size for the peak hour or set a clear queueing policy.
Typical tools
| Layer | Common choices |
|---|---|
| Model servers | vLLM, NVIDIA Triton with TensorRT-LLM, Text Generation Inference |
| Platform | Kubernetes with NVIDIA GPU Operator, KServe |
| Gateway | An API gateway with authentication and per-team quotas |
| Monitoring | Prometheus and Grafana, DCGM, request tracing |
Connecting a model to your own documents? See RAG solutions. Choosing the model itself? See enterprise LLM.
More AI infrastructure guides: Training cluster · GPU network fabric · Storage for AI · Power and cooling · FAQ · Use case: private LLM platform
Planning a GPU cluster?
Tell us the models you want to run or train, how many users, and where it will be hosted. We will come back with a first sizing: GPUs, servers, network, storage, and the power and cooling your data centre will need.