Home / AI infrastructure / Inference platform
AI infrastructure guide

GPU inference platform

Inference is where a model meets its users, and where most of the GPU budget ends up over time. The goals are different from training: steady low latency, many users at once, and as many requests per GPU as possible without making anyone wait.

Measure inTokens per second, latency
Model serversvLLM, Triton, TensorRT-LLM
Share GPUs withMIG, time-slicing
Scales onQueue depth
Reference architecture

How requests are served

Scroll sideways to see the whole diagram →
Serving layerPlatformUsers and applicationschat, search, internal toolsAPI gatewaysign-in, rate limits, usage per teamModel routersends each request to the right modelModel servers70B on 2-4 GPUs, 8B on a GPU sliceAutoscaleradds replicas when queues growMonitoringlatency, tokens per second, cost per teamModel registryapproved, versioned modelsGPU serversH100 / H200 for large, L40S for smallKubernetes with GPU Operatordrivers, GPU sharing, scheduling1234
Numbered lines show the main flows from a user request to a GPU and back.
PartWhat it does
API gatewayEvery request is signed in, rate-limited and counted per team, so cost can be charged back and one team cannot starve another.
1 RoutingSimple questions go to a small, cheap model. Hard ones go to a large model. This alone can cut GPU needs sharply.
2 Model registryOnly approved, versioned models are served. Rolling back a bad model is a configuration change, not an emergency.
3 AutoscalingReplicas are added when request queues grow and removed when they shrink, within the GPUs available.
4 MonitoringTime to first token, tokens per second, GPU utilisation and cost per team, on one dashboard.

Choosing GPUs for inference

Model sizeFits onNotes
~8BOne L40S (48 GB) or a slice of an H100Cheapest to serve. Often good enough for classification, extraction and simple chat.
~70B2 to 4 x H100 80 GB at 16-bit; 1 to 2 x H100 at 8-bit; 1 x H200 at 8-bitThe sweet spot for many enterprise assistants. 8-bit precision usually costs little quality.
~400BOne full 8-GPU H200 server at 8-bit, or moreRarely needed for enterprise use. Consider whether a 70B model with good retrieval does the job.

Memory is only the start. Many concurrent users with long contexts need extra memory for the KV cache. Use the GPU calculator to see the effect.

Getting it right

  • Batch requests. Model servers such as vLLM serve many users on one GPU by batching continuously. Serving one request at a time wastes most of the GPU.
  • Share GPUs for small models. MIG splits an H100 into up to seven isolated slices. Small models do not need a whole GPU each.
  • Quantise with care. 8-bit weights roughly halve memory with little quality loss for most models. Test 4-bit on your own tasks before relying on it.
  • Set latency targets per use case. A chat assistant needs a fast first token. An overnight document job does not. Size for the strictest case only where it applies.
  • Plan capacity for peaks. Usage often doubles at the start of the working day. Size for the peak hour or set a clear queueing policy.

Typical tools

LayerCommon choices
Model serversvLLM, NVIDIA Triton with TensorRT-LLM, Text Generation Inference
PlatformKubernetes with NVIDIA GPU Operator, KServe
GatewayAn API gateway with authentication and per-team quotas
MonitoringPrometheus and Grafana, DCGM, request tracing

Connecting a model to your own documents? See RAG solutions. Choosing the model itself? See enterprise LLM.

Planning a GPU cluster?

Tell us the models you want to run or train, how many users, and where it will be hosted. We will come back with a first sizing: GPUs, servers, network, storage, and the power and cooling your data centre will need.