Home / Enterprise LLM / Private hosting
LLM guide

Private LLM hosting

Running an open model inside your own network means prompts, documents and answers never leave infrastructure you control. It also means you run the platform. Here is what a sound private LLM platform includes beyond the model itself.

DataStays in your network
Model serversvLLM, Triton
Must haveGateway, guardrails, logs
Can beFully air-gapped
Reference architecture

How the pieces fit together

Scroll sideways to see the whole diagram →
Your applications and controlsLLM platform (inside your network)Staff and systemschat, documents, business appsAssistants and applicationscall the LLM through one APIAudit logs and security monitoringwho asked what, and whenIdentity providergroups decide who may use which modelApproved data sourcesused through RAG, not trainingLLM gatewaysign-in, quotas, routing, loggingGuardrailsmask personal data, filter contentModel serversopen models on your GPUsModel registry and test resultsonly approved versions are served12345
Every request passes through the gateway and guardrails before it reaches a model, and every request is logged.
PartWhat it does
1 One entry pointApplications never call a model directly. They call the gateway, which checks who is asking and how much they may use.
2 GuardrailsPersonal data such as Aadhaar or account numbers can be masked before the model sees it, and harmful content filtered on the way out.
3 Model serversOpen models served on your GPUs. See the inference platform guide for sizing.
4 Audit trailRequests and decisions are logged to your security monitoring, so you can answer "who asked what" later.
5 Approved models onlyA new model or version is served only after it passes your test set.

Air-gapped deployments

  • Models, container images and updates are brought in through a controlled, scanned transfer process.
  • No component needs internet access at runtime: no telemetry, no licence checks against external servers.
  • Plan how updates reach the platform before it is built. Air-gapped systems that cannot be updated become insecure.

Air-gapping suits defence, some government and some financial workloads. For most organisations, a private network with strict outbound rules is enough.

Typical tools

LayerCommon choices
Model serversvLLM, NVIDIA Triton with TensorRT-LLM, Text Generation Inference
GatewayLiteLLM or an API gateway with LLM routing, quotas and logging
GuardrailsMicrosoft Presidio for personal data, NVIDIA NeMo Guardrails, custom rules
PlatformKubernetes with NVIDIA GPU Operator
ObservabilityPrometheus, Grafana, Langfuse or OpenTelemetry traces

Starting with LLMs, or stuck after a pilot?

Tell us the task you want AI to help with and any rules about where your data can go. We will come back with a plain recommendation: which kind of model, where to run it, roughly what it costs, and how to know if it is working.