Running an open model inside your own network means prompts, documents and answers never leave infrastructure you control. It also means you run the platform. Here is what a sound private LLM platform includes beyond the model itself.
DataStays in your network
Model serversvLLM, Triton
Must haveGateway, guardrails, logs
Can beFully air-gapped
Reference architecture
How the pieces fit together
Scroll sideways to see the whole diagram →Every request passes through the gateway and guardrails before it reaches a model, and every request is logged.
Part
What it does
1 One entry point
Applications never call a model directly. They call the gateway, which checks who is asking and how much they may use.
2 Guardrails
Personal data such as Aadhaar or account numbers can be masked before the model sees it, and harmful content filtered on the way out.
Requests and decisions are logged to your security monitoring, so you can answer "who asked what" later.
5 Approved models only
A new model or version is served only after it passes your test set.
Air-gapped deployments
Models, container images and updates are brought in through a controlled, scanned transfer process.
No component needs internet access at runtime: no telemetry, no licence checks against external servers.
Plan how updates reach the platform before it is built. Air-gapped systems that cannot be updated become insecure.
Air-gapping suits defence, some government and some financial workloads. For most organisations, a private network with strict outbound rules is enough.
Typical tools
Layer
Common choices
Model servers
vLLM, NVIDIA Triton with TensorRT-LLM, Text Generation Inference
Gateway
LiteLLM or an API gateway with LLM routing, quotas and logging
Guardrails
Microsoft Presidio for personal data, NVIDIA NeMo Guardrails, custom rules
Platform
Kubernetes with NVIDIA GPU Operator
Observability
Prometheus, Grafana, Langfuse or OpenTelemetry traces
Tell us the task you want AI to help with and any rules about where your data can go. We will come back with a plain recommendation: which kind of model, where to run it, roughly what it costs, and how to know if it is working.