Home / Whitepapers / Enterprise LLM
Whitepaper · Enterprise LLM

Running LLMs in the Enterprise

Choosing between API and self-hosted models, governing them, and keeping costs predictable. Written for CIOs, IT heads, enterprise and AI architects, and security, risk and data protection teams.

Reading time26 minutes
LengthApprox. 15 pages
Editionv1.0, October 2026
FormatWeb and PDF

What is inside

  • A simple way to triage use cases by value, risk and how easily answers can be checked
  • Public API, private endpoint, self-hosted and hybrid compared on data, quality and cost
  • A cost crossover formula: the daily token volume at which your own GPUs pay off
  • How to choose and test models on your own work, not on leaderboards
  • When to prompt, when to use RAG and when fine-tuning is worth it
  • An LLM gateway design for routing, PII redaction, budgets and audit, with DPDP Act points
  • A 36-point readiness checklist for platform, security and operations
Or start reading online ↓
Download the full PDFApprox. 15 pages, including the 36-point checklist

We confirm your email with a one-time code, then the PDF downloads straight away. No newsletters unless you ask.

01

Executive summary

Most organisations already use large language models. Often the first use is unofficial: staff pasting documents into free chat tools on personal accounts. The question for IT leadership is no longer whether to use LLMs, but how: which tasks, which models, where they run, what data they may see, and how anyone knows the answers are good enough and the bill is under control.

This paper sets out a practical approach for the IT leaders who decide and the security, risk and data protection teams who sign off.

Five things to take away

  1. Start from tasks, not models. Pick a few frequent, well-defined tasks where answers can be checked. A narrow task done well beats “AI everywhere” done badly.
  2. Test on your own examples. A test set of 100 to 300 real cases, with answers agreed by the people who do the work, is the single most valuable asset in an LLM programme.
  3. Self-hosting is a data and volume decision. It pays off on cost only above a daily token volume you can calculate. Below that, data rules are the only good reason.
  4. Put one gateway in front of every model. Sign-in, redaction, routing, budgets and logging in one place, so models can change without changing every application.
  5. Treat models like any other production change. New models and prompts go through test, approval and gradual rollout, with a quick way back.
02

Why LLM projects stall

LLM pilots are easy to start and hard to finish. A demo on ten hand-picked questions impresses in a week; handling thousands of messy, real requests safely and at a known cost takes discipline most pilots skip. The same problems recur:

What we findWhat happens
No approved option for staffPeople use whatever is free. Customer data and internal documents leave the company without anyone deciding they should.
Pilot judged on easy questionsReal users ask vague, multi-part and out-of-scope questions. Quality drops, and nobody measured it before go-live.
Every app calls the model directlyKeys are scattered in code, nobody can see total usage, and changing the model means changing every application.
Long prompts on the largest modelWhole documents are sent with every request to the most expensive model. Monthly cost becomes a board-level question.
Fine-tuning to teach factsMonths go into training a model to “know” policies that change every quarter, when the real need was to look them up.
No owner after go-liveQuality drifts after a model update or a document change, and complaints are the first warning.
In short: most LLM failures are not model failures. They come from unclear tasks, untested quality, uncontrolled access to data and unmanaged cost.
03

The building blocks, explained properly

Tokens and context

Models read and write in tokens, pieces of words. In English one token is roughly three quarters of a word, so 1,000 words is about 1,300 tokens. Hindi and other Indian languages in their own scripts often take two to four times as many tokens for the same meaning on many models, which raises cost and fills the context faster. The context window is how many tokens the model considers at once. Output tokens usually cost three to five times more than input.

Model size

Small open models of around 8 billion parameters (8B) are fast and cheap, and handle extraction, classification and routing well. Models around 70B handle most drafting, summarising and question answering. The largest commercial models lead on complex reasoning and long documents.

Serving

An open model runs on a model server such as vLLM or NVIDIA Triton with TensorRT-LLM. Three techniques let one server handle hundreds of users rather than a few:

  • Continuous batching: requests join and leave the batch at every step, so the GPU stays busy instead of waiting for the slowest request.
  • Paged KV cache: the memory that holds each conversation’s context is allocated in small blocks, so many more requests fit on one GPU.
  • Prefix caching: a long system prompt or a shared document is processed once and reused across requests.

Two numbers describe the user experience: time to first token (how long before the answer starts, ideally under one to two seconds for chat) and output tokens per second per user (reading speed is about 10 to 15; 30 or more feels instant).

Quantisation

Quantisation stores weights in fewer bits. FP8 halves memory against 16-bit with little measurable quality loss on most tasks, and runs natively on H100, H200 and L40S GPUs. 4-bit formats such as AWQ halve it again but can hurt accuracy, especially in Indian languages. Re-run the test set on any quantised model.

Prompting, RAG or fine-tuning

These are often confused. They solve different problems and are usually used together:

The problemTry firstWhy
The model does not know our policies, products or casesRAG: look up the relevant passages and give them to the modelFacts change. Looking them up at question time keeps answers current and lets you show sources.
Answers are in the wrong format or toneBetter prompts with two or three good examples, then fine-tuningExamples in the prompt fix most format problems. Fine-tune only if it still drifts.
A narrow task repeated thousands of times a dayFine-tune a small modelA small fine-tuned model can match a large general one on one task at a fraction of the cost.
Specialised language, such as legal or internal jargonRAG plus a glossary, then fine-tuning if neededRetrieval supplies the facts; fine-tuning helps the model use the terms naturally.

Fine-tuning today usually means LoRA: training a small set of extra weights rather than the whole model. It needs hundreds to a few thousand good examples, fits on one 8-GPU server for a 70B model, and usually has to be repeated when the base model changes.

04

API, private endpoint, self-hosted or hybrid

There are four ways to give the organisation access to LLMs. They differ in where data goes, which models are available, how cost behaves as usage grows, and how much the IT team has to run.

OptionWhere data goesModelsCost behaviourWhat you run
Public APITo the provider, under its enterprise termsStrongest commercial modelsPure per-token; grows in line with useAlmost nothing
Private cloud endpointStays in your cloud tenancy and chosen regionCommercial models offered in that regionPer-token or reserved capacityCloud networking, keys, access
Self-hosted open modelNever leaves your networkOpen-weight models (Llama, Mistral, Qwen, Gemma and others)Mostly fixed; cheap per token at high useGPUs, model servers, platform, upgrades
Hybrid behind a gatewayDecided per request by data classBothFixed base plus per-token for the restGateway and a smaller GPU estate

Public API

The fastest start and the strongest models. The questions are contractual: whether prompts are retained or used for training, where they are processed, and what cost looks like at ten times today’s volume. Enterprise agreements usually exclude training on your data; consumer terms often do not.

Private cloud endpoint

Commercial models deployed inside your own cloud account, in an Indian region where available. Data stays within your tenancy, under controls you already use. Check which models each region offers.

Self-hosted open model

An open-weight model on your own GPUs. Data never leaves infrastructure you control, versions change only when you decide, and it can run air-gapped. In exchange you run GPUs, serving and upgrades, and open models may trail the best commercial ones on the hardest tasks. Read each licence: some restrict use above a user count.

Hybrid behind a gateway

Most organisations end up here: confidential and personal data goes to self-hosted models, while redacted public and internal content may use a commercial model. A gateway makes the choice per request.

Taking this to a meeting? Get the PDF to share with your team: same content, printable checklist, no clutter.
Download the PDF
05

Choosing: use cases first, then models

Triage the use cases

Score each candidate on frequency, value, data sensitivity and how easily a person can check the output. The best first projects are frequent, valuable and easy to check.

Use case typeExamplesRiskTypical first approach
Draft for a person to editEmail replies, meeting notes, first drafts of reportsLow: a person reviewsAPI or private endpoint; light controls
Answer from documentsPolicy and product questions, IT and HR helpMedium: wrong or leaked answersRAG with permissions and sources
Extract and classifyClaim forms, invoices, tickets, KYC documentsMedium: errors flow into systemsSmall or fine-tuned model, confidence thresholds
Customer-facing conversationService chat, contact centre assistHigh: brand and conduct riskAssist an agent first, then limited self-service
Decisions about peopleCredit, claims, hiring, eligibilityVery high: fairness and regulationDecision support only; a person decides

Five questions that decide hosting

  1. May this data leave our network? Classify the data each use case touches. If personal, customer financial or restricted government data is involved, sector rules and your own policy may decide for you.
  2. What is the daily token volume, now and in a year? Compare it with the crossover point in the sizing section. Below it, APIs are cheaper.
  3. Does an open model pass the test set? If only a frontier commercial model meets the quality bar, self-hosting the task is not yet an option.
  4. Is there a hard latency, availability or air-gap requirement? Plants, defence and some government networks cannot depend on an internet path.
  5. Can we run GPUs well? If no team owns drivers, serving and capacity, budget for one or use a managed private option.

Testing models on your own work

Collect 100 to 300 real examples with the answer a good employee would give, including messy and out-of-scope cases, where models differ most. Define “good”: correct facts, right format, and “I don’t know” when it should. Score two or three candidates automatically and by reviewers, then weigh cost and speed. A model 3 percent better but ten times dearer is rarely right. Keep the test set: a new model version can then be judged in an afternoon.

06

Reference architecture

The diagram shows a hybrid platform built around one LLM gateway. Applications never call a model directly: the gateway knows who is asking, cleans the request, picks a model, enforces budgets and records what happened.

APPLICATIONS AND USERSLLM GATEWAY, IN YOUR NETWORKOBSERVABILITY AND GOVERNANCESELF-HOSTED MODELS ON GPUSAPPROVED EXTERNAL APIStaff assistantbrowser and TeamsBusiness appsclaims, CRM, ITSMRAG servicepolicy documentsDevelopersone key per teamGateway entry pointcompany sign-in, team keysBudgets and limitsper team, per monthPII redactionAadhaar, PAN, account numbersGuardrailsinjection and content filtersRouterby data class, task and costModel catalogueapproved versions onlyLarge model (70B)vLLM, FP8, 4 GPUsSmall model (8B)extraction and routingEmbedding modelfor RAG searchCommercial modelPublic and Internal onlyAudit log and tracesretention set by policyEvaluation and costtest set, dashboards4Confidential stays in house5redacted first6logged123User or API trafficControl / API callData / replicationLogging / management

Numbered flows: (1) every application and user calls one gateway endpoint, authenticated by company sign-in or a team key, (2) each request passes PII redaction and guardrails, (3) the router reads the data class tag and task type, (4) Confidential and Restricted requests go only to self-hosted models, with the small model handling extraction and routing, (5) Public and Internal requests may go to an approved commercial model after redaction, (6) every request, model choice and cost is logged for audit and evaluation.

What the gateway does

FunctionIn practice
Identity and accessCompany sign-in for people, one key per team for applications. Groups decide which models and data classes each team may use.
RedactionPersonal data is masked before any model sees it: names, phone numbers, Aadhaar, PAN, bank account and IFSC codes, using Microsoft Presidio plus rules for Indian identifiers.
RoutingBy data class, task and cost: small model for extraction, larger model for drafting, commercial model only for permitted classes. Fallback to a second model if one is down.
Budgets and limitsMonthly budget per team with alerts at 50, 80 and 100 percent, a cap on tokens per request, and rate limits so one runaway job cannot starve everyone else.
LoggingWho asked, which model answered, tokens, cost, latency and guardrail decisions. Prompt text is kept only as long as policy allows.
Common tools: LiteLLM or an API gateway with LLM routing, vLLM on Kubernetes with the NVIDIA GPU Operator, Presidio and NeMo Guardrails, and Langfuse or OpenTelemetry for traces. All are replaceable; the pattern is what matters.
07

Sizing and cost

The cost crossover

An API costs a fixed amount per token. A self-hosted model costs roughly the same each month whether it is busy or idle. The volume at which the two are equal is the crossover:

Crossover tokens per day = monthly self-hosted cost (₹) × 1,000,000 ÷ (active days per month × blended API price per million tokens (₹))

The blended price mixes input and output at your real ratio. For a frontier-class model at ₹250 per million input tokens and ₹1,250 per million output, a 3:1 input to output mix blends to (3 × 250 + 1,250) ÷ 4 = ₹500 per million. Mid-tier commercial models are often a quarter of that or less.

What self-hosting costs per month

Planning ranges for one 8-GPU H100 server running a 70B-class model in FP8 as two 4-GPU replicas:

ItemBasisMonthly, planning range
Server, networking and storage share₹2.6 to 3.2 crore, over 48 months₹5.4 to 6.7 lakh
Power and coolingAbout 10.2 kW × 1.5 PUE × 730 hours = 11,200 kWh at ₹8 to 10₹0.9 to 1.1 lakh
Rack space in a data centre or colocationOne high-density rack position₹0.8 to 1.5 lakh
Support, software and sparesHardware support, enterprise software where used₹0.5 to 1.0 lakh
Platform engineeringAbout 1.5 people across GPUs, serving and gateway₹2.5 to 3.5 lakh
Total₹10 to 14 lakh, midpoint ₹12 lakh

With ₹12 lakh a month, 22 working days and a frontier-class API at ₹500 per million tokens, the crossover is 12,00,000 × 1,000,000 ÷ (22 × 500), or about 110 million tokens a day. Against a mid-tier API at ₹120 per million it rises to about 450 million tokens a day, which is more than one server can handle. Self-hosting wins on cost only when the task needs a large model and volume is high.

GPU memory

GPU memory needed ≈ model weights + (KV cache per token × context length × concurrent requests) + 10 to 20% overhead

A 70B model’s weights need about 140 GB in 16-bit, 70 GB in FP8 and 35 to 40 GB in 4-bit; its KV cache is about 0.33 MB per token. On four 80 GB H100s (320 GB), FP8 weights take 70 GB and overhead about 40 GB, leaving roughly 210 GB for cache: about 640,000 tokens, or around 80 concurrent requests at 8,000 tokens of context each.

A worked example

A financial services firm plans two uses. A staff assistant: 1,500 daily users × 12 questions × 4,000 tokens (3,500 in, 500 out) = 72 million tokens a day. Loan file summaries: 1,000 files × 30,000 tokens (28,000 in, 2,000 out) = 30 million a day. Total about 102 million tokens a day, or 2,244 million a month over 22 working days.

Frontier APIMid-tier APISelf-hosted 70B, one server
Monthly cost at 102 M tokens a day2,244 × ₹500 = ₹11.2 lakh2,244 × ₹120 = ₹2.7 lakh₹10 to 14 lakh
Monthly cost at 3 times the volume₹33.7 lakh₹8.1 lakh₹20 to 24 lakh, with a second server
Passes the test set?YesFor the assistant, not for summariesYes, after prompt tuning
Customer data leaves the network?Yes, under contractYes, under contractNo

Peak load check: 11 million output tokens over a 9-hour day, doubled for the busy hour, is about 680 tokens a second, within one server’s planning capacity of around 2,000. The sensible answer is hybrid: loan files hold customer financial data, so they run self-hosted; the assistant uses the mid-tier API on redacted internal content, and moves in-house if volume grows.

Costs that are often missed: long prompts repeated on every call, evaluation runs, a fallback for availability, GPU refresh after four to five years, and the people to run it.
08

Security, data protection and the DPDP Act

Data protection in India

The Digital Personal Data Protection Act, 2023 and its Rules, whose main obligations phase in through 2026 and 2027, apply to personal data in prompts, documents and logs just as to any other system. In practice:

  • Purpose and notice: personal data in an LLM workflow needs a lawful basis and a stated purpose. Training a model on customer data collected for servicing is a new purpose.
  • Minimisation: redact personal data the task does not need before it reaches any model, internal or external.
  • Retention and erasure: prompts and traces containing personal data are personal data. Set retention periods; make erasure reach logs and indexes.
  • Security safeguards and breach reporting: reasonable safeguards are a legal duty, and breaches must be reported to the Data Protection Board and affected people. CERT-In expects cyber incidents to be reported within 6 hours.
  • Processors: an API provider processing personal data on your behalf needs a contract that matches your obligations.
  • Sector rules: RBI, IRDAI and SEBI requirements on outsourcing, data location and audit may be stricter than the Act. Check them per use case.

Threats specific to LLMs

ThreatControl
Prompt injection: hidden instructions in a document or emailTreat retrieved text as data; limit what the model can do; filter outputs; test with attack examples.
Leakage across usersPermissions enforced at retrieval time; no shared conversation memory between users.
Sensitive data in logsRedact before logging; restrict who can read traces; time-limited retention.
Unapproved models and keysAll traffic through the gateway; outbound access to AI services blocked except via the gateway.
Untrusted model filesDownload from known sources, verify checksums, prefer safetensors format, scan containers.
A useful reference: the OWASP Top 10 for LLM Applications lists prompt injection first. Use it as the basis of the security review before go-live.
09

Operating it: evaluation, monitoring and change

An LLM system that worked in March can be worse in June without anyone touching it: the provider updates the model, documents change, users ask new things.

Measure quality all the time

  • Offline: re-run the test set before every change to model, prompt, retrieval or redaction rules. No change goes live with a lower score without a written reason.
  • Online: thumbs up or down with an optional comment, plus a weekly sample of 30 to 50 live answers reviewed by people who know the work.
  • Drift: alert on falling scores, rising complaints, more refusals or a change in answer length after any change.

Watch the platform

MetricWhy it mattersTypical alert
Time to first token, 95th percentileUsers feel this firstAbove 2 seconds for chat
GPU memory and cache useFull cache means queued requestsAbove 90% for 15 minutes
Queue length and rejected requestsCapacity is shortAny sustained queue at busy hour
Tokens and cost per teamBudgets and surprises80% of monthly budget
Error and fallback rateA model or provider is failingAbove 1% of requests

Change models safely

Keep a catalogue of approved models with version, licence, allowed data classes and test score. A new version passes the test set and a security check, then takes a small share of gateway traffic before full rollout, with a one-step rollback. Pin API model versions where possible.

10

A practical roadmap

A first production use case typically takes 10 to 14 weeks.

StageTypical durationOutcome
1. Triage2 weeksUse cases scored, one or two chosen with business owners, data classified, baseline time and cost measured.
2. Test and compare3 to 4 weeksTest set of 100 to 300 real cases, two or three models scored on quality, speed and cost, hosting decision.
3. Guarded build3 to 5 weeksGateway with sign-in, redaction, routing, budgets and logging; GPU platform if self-hosting; security review.
4. Pilot4 weeksReal users in one department, quality and time saved measured against baseline, issues fixed.
5. Scale and operateOngoingMore use cases on the same gateway, monthly quality review, model catalogue and change process in use.
Order matters: build the gateway with the first use case, not after the fifth. Retrofitting it later is slow and rarely complete.
11

Ten common mistakes

  1. No approved option for staff. Shadow use of public tools continues with company data.
  2. Choosing a model from a leaderboard. Test on your own examples instead.
  3. Testing only on easy questions. Real users find the edge cases on day one.
  4. Fine-tuning to teach facts. Use RAG for anything that changes.
  5. Self-hosting below the crossover without a data reason. GPUs sit idle and cost more than the API would.
  6. Applications calling models directly. No central view of usage, cost or data flows.
  7. Sending whole documents on every call. Retrieve the relevant passages and cache shared prompts.
  8. Logging everything forever. Logs become the largest store of personal data in the company.
  9. Letting the provider change the model silently. Pin versions and re-test before switching.
  10. No owner after go-live. Quality drifts and nobody notices until users stop using it.
12

Enterprise LLM readiness checklist (36 points)

Use this list to score your current position. Anything you cannot tick with evidence is a gap worth closing.

Use cases and ownership

  • An approved LLM option available to staff, with a usage policy
  • Use cases scored on frequency, value, data sensitivity and ease of checking
  • A named business owner for each use case in production
  • Baseline time and cost measured before the pilot
  • Decisions about people kept with people, with the model as support only
  • A shared register of LLM use cases, models and data classes
Data and compliance 7 points Models and evaluation 6 points Platform and serving 5 points Gateway, cost and security 7 points Operations and change 5 points

The remaining 30 points are in the PDF, laid out as a printable checklist.

Get the full checklist
13

Glossary

TermMeaning
TokenA piece of a word. Models read, write and are priced in tokens; about 1,300 tokens per 1,000 English words.
Time to first tokenHow long a user waits before the answer starts to appear.
Continuous batchingServing technique that adds and removes requests at every step to keep the GPU busy.
KV cacheGPU memory holding the processed context of each active request.
QuantisationStoring model weights in fewer bits (FP8, 4-bit) to save memory and speed up serving.
RAGRetrieval-augmented generation: finding relevant passages and giving them to the model to answer from.
LoRAA fine-tuning method that trains a small set of extra weights instead of the whole model.
LLM gatewayA single entry point for all model traffic that handles access, redaction, routing, budgets and logging.
Prompt injectionInstructions hidden in text the model reads, intended to make it ignore its rules.
Data FiduciaryUnder the DPDP Act, the organisation that decides why and how personal data is processed.

About Vakratron Systems

Vakratron Systems is a vendor-neutral infrastructure design firm. We design data centre, disaster recovery, cloud, GPU and AI platforms for enterprises and government buyers, write our assumptions down, and stay with a design until it is running and tested.