One LLM gateway for thirty teams, and customer data never leaves the building
A large non-banking lender found that more than thirty teams wanted to use language models, and several had already started on their own with personal API keys. This design replaces that sprawl with one gateway: every request is classified, cleaned of personal data, routed to the right model by sensitivity, logged and charged to the team that made it.
Thirty teams, thirty ways of using AI, and no one could say where the data went
The lender serves about 4 million retail and small-business borrowers through 600 branches and a digital channel. Over twelve months, teams in collections, customer service, credit, engineering, risk and HR had each found their own uses for language models: drafting credit notes, summarising call transcripts, answering policy questions, writing code.
A security review found 14 different external AI services in use, several paid on corporate cards, and at least two cases where loan application text containing names, PAN and phone numbers had been pasted into a public chatbot. Nobody could produce a list of which prompts had left the company, or what they cost.
The CIO did not want to stop the work: some of it was saving real time. The ask was for one controlled way in, where the safe choice is also the easy one for teams, customer data stays on company hardware, and the board can see usage, cost and risk on a single page.
What could not be compromised
- Customer personal data (KYC, loan, repayment, contact details) must never be sent to a model outside the company’s data centre.
- DPDP Act obligations: purpose limitation, data minimisation, access control and the ability to show what personal data was processed and why.
- RBI expectations on IT governance and outsourcing: external AI providers are treated as material vendors, with contracts, audits and exit plans.
- Teams must not need to rewrite applications every time a model changes. One API, one set of credentials per team.
- Spend on external APIs is capped per team and visible to finance every month.
Four ways to give teams access to models
Each option was scored on data control, cost at 250,000 requests a day, speed for teams, and the effort to run it with a platform team of six.
Every request takes the same path, and the data class decides where it can go
Applications call one endpoint with a team key. The gateway checks the team’s quota, classifies the request, removes personal data where needed, and only then picks a model. Confidential and Restricted data can only reach models on the company’s own GPUs. The network enforces the same rule: only the gateway has a route to external AI providers.
| Building block | Why it is there |
|---|---|
| 1 Gateway API | One OpenAI-compatible endpoint for every team. Staff sign in through company SSO; applications use a team key tied to a named use case. Teams switch models by changing a name in the request, not their code. |
| 2 Classify and redact PII | A small classifier tags each request as Public, Internal, Confidential or Restricted, using the declared use case plus what it finds in the text. Pattern rules and an NER model detect Aadhaar, PAN, account and loan numbers, phone numbers, emails and names, and replace them with tokens before anything is logged or routed. If detection and declaration disagree, the stricter class wins. |
| 3 Policy router | Chooses the model from the catalogue by data class first, then by cost and latency. Restricted and Confidential traffic has no route to external providers at all. If the self-hosted pool is busy, requests queue; they never spill outside. |
| 4 Self-hosted models | A 70B-class open model on 8 x H100 for harder reasoning and long documents, an 8B-class model on L40S cards for summaries and classification, and embedding models for teams building RAG search. Served with continuous batching. |
| 5 Approved external APIs | Two contracted providers with zero data retention, no training on prompts, and audit rights. Reached only from the gateway’s egress proxy; the firewall blocks AI provider domains from every other subnet. |
| 6 Budgets, rate limits and cache | Each team gets a monthly token budget and requests-per-minute limit per use case. Alerts at 80%, hard stop at 100% unless the budget owner raises it. Exact and semantic caching answer repeated policy and FAQ questions without calling a model. |
| 7 Prompt and response log | Every request is stored after redaction with team, use case, model, tokens, cost and latency. Raw text is kept only for Restricted use cases that need it for audit, encrypted and readable by two named roles. |
| 8 Model catalogue and evaluation | A register of approved models, what data classes each may receive, owners and review dates. New models and prompt changes are scored against each team’s test set before they go live. |
GPU capacity sized from request mix, not from model size alone
Request volumes come from three months of logs on the pilot teams, scaled to all 30 teams. About 70% of requests carry Confidential or Restricted data, so most traffic has to stay on company GPUs.
| Item | Figure | Basis |
|---|---|---|
| Requests | ~250,000 a day, peak ~15 a second | 30 teams, 10-hour working day, peak hour at 2x average |
| Tokens per request | ~2,500 in, ~350 out | Measured in pilot; RAG and document tasks pull the average up |
| Self-hosted share | ~70%, ~10.5 requests a second at peak | Confidential and Restricted share from pilot classification |
| Large model | 8 x H100 (2 replicas x 4 GPUs) | ~40% of self-hosted traffic = ~4.2 req/s x 350 = ~1,500 output tokens/s at FP8; one replica handles it, two for failover |
| Small model and embeddings | 4 x L40S | ~6.3 req/s on the 8B model (2 GPUs) plus embeddings and the PII model (2 GPUs) |
| External API spend | ₹8 to 12 lakh a month | ~30% of requests, ~4.7 billion tokens a month at blended contract rates, split into team caps |
| Log storage | ~1.2 TB a year | ~250,000 requests x ~13 KB compressed x 365 days, kept per retention policy |
Before the H100 order, the chosen 70B model is benchmarked on rented capacity with replayed pilot traffic. The cache is expected to absorb 15 to 20% of requests once policy and FAQ use cases are live, which is held as headroom rather than used to cut GPUs.
From two pilot teams to thirty, with the side doors closed last
The gateway has to be the easiest way in before the other ways are closed. Teams move in waves, each with a named use case and a test set.
Policy and design
Weeks 1 to 4
Data classes agreed with risk, compliance and the DPO. Use-case register started. External providers assessed as material vendors.
Gate: Classification policy and provider contracts signed off.
Build the gateway
Weeks 5 to 10
Gateway, redaction, router, logging and budgets built. GPU servers installed and models benchmarked.
Gate: Redaction recall above 99% on a labelled set of 5,000 real prompts.
Pilot teams
Weeks 11 to 14
Customer service and engineering move first. Shadow classification compares declared and detected data classes.
Gate: No Restricted request routed externally in four weeks of logs.
Onboard all teams
Weeks 15 to 18
Remaining teams join in three waves, each with budgets, keys and an evaluation set.
Gate: Every active AI use case listed in the register with an owner.
Close side doors
Weeks 19 to 20
Firewall blocks AI provider domains except from the gateway. Corporate-card AI subscriptions cancelled.
Gate: Network logs show no direct AI traffic for two weeks.
The risks that matter with a shared AI gateway
| Risk | What could happen | How the design handles it |
|---|---|---|
| Missed personal data | Redaction misses a name or account number in free text | Stricter class wins on disagreement, so a missed item still goes to a self-hosted model. Weekly sampled review of logs; recall measured against a growing labelled set. |
| Gateway as single point of failure | An outage takes AI out of 30 teams at once | Three gateway instances across two racks, stateless behind a load balancer. Applications must handle AI being unavailable. |
| Cost runaway | A looping script burns a month’s budget overnight | Per-minute rate limits, hard monthly caps and alerts at 80%. External spend is reported by team to finance every month. |
| Prompt injection through documents | A customer document tries to change a model’s instructions | Document text is passed as data, system prompts are fixed per use case, and the gateway gives models no tools by default. |
| Vendor change or exit | An external provider changes terms or retention | Two approved providers behind one API. Contracts include exit clauses; the catalogue lets traffic move in a day. |
What was optimised
Kept on company GPUs
All Confidential and Restricted traffic runs on self-hosted models. Only redacted, lower-risk work goes outside.
Endpoint for every team
One API and one key per use case. Swapping a model is a catalogue change, not a code change.
Answered from cache
Repeated policy and FAQ questions are served without calling a model.
Right model for the job
Summaries and classification go to the small model. The 70B model is kept for work that needs it.
Budget alerts before the cap
Team owners see spend daily and get a warning well before a hard stop.
Direct routes to AI providers
The firewall makes the gateway the only way out, so policy cannot be bypassed.
What the design is built to deliver
| Measure | Before | Design target |
|---|---|---|
| External AI services in use | 14, unmanaged | 2 approved providers, behind the gateway |
| Customer data sent to external models | Unknown | None, enforced by routing and network |
| AI spend visibility | Scattered card payments | Monthly report by team and use case |
| Time to give a new team access | Weeks of ad hoc approval | 2 to 3 days with a registered use case |
| Cost per 1,000 requests | Not measured | 25 to 40% lower than pilot-era external spend for the same work |
| Audit evidence for DPDP and RBI reviews | Manual interviews | Use-case register, data-class logs and provider contracts in one pack |
Targets are confirmed during the pilot on real traffic. GPU sizing and external spend depend on the final model choices and on how many use cases turn out to need the large model.
What a team needs to deliver this
LLM platform design
Gateways, routing, catalogues and APIs that let many teams use models safely through one door.
Data classification and PII redaction
Detection of Indian identifiers, NER tuning and recall testing on real prompts.
GPU serving and capacity planning
Model selection, quantisation, batching and sizing from measured request mix.
AI governance
Use-case registers, model approval and evaluation that satisfy risk, compliance and the DPO.
Regulatory mapping
Turning DPDP Act and RBI outsourcing expectations into controls an auditor can test.
Network and egress control
Proxy and firewall design so that policy is enforced by the network, not by goodwill.
Other scenarios
Agentic AI for accounts payable
An AI agent reads, checks and posts supplier invoices for a manufacturing group, with people approving every exception.
Read →Contact centre AIAI copilot for a service contact centre
A consumer-durables brand gives 600 service agents a self-hosted AI copilot: live transcripts, answers from manuals, automatic call notes and quality scores on every call.
Read →MLOpsMLOps platform for demand forecasting
An FMCG company turns laptop notebooks into an MLOps platform that trains, approves, serves and retrains 2,000-SKU demand forecasts on its own.
Read →Do you know how many AI services your teams are already using?
Share a rough list of AI use cases and the data they touch. We will come back with a plain view of what can safely go to external models, what must stay on your own hardware, and what a gateway would need to look like.