Home / Deployment scenarios / Enterprise LLM gateway for a lender
Reference deployment Enterprise LLM platform

One LLM gateway for thirty teams, and customer data never leaves the building

A large non-banking lender found that more than thirty teams wanted to use language models, and several had already started on their own with personal API keys. This design replaces that sprawl with one gateway: every request is classified, cleaned of personal data, routed to the right model by sensitivity, logged and charged to the team that made it.

SectorNBFC, retail and MSME lending
Users30+ teams, ~250,000 requests/day
Programme20 weeks to all teams
ModelHybrid: self-hosted + approved APIs
The situation

Thirty teams, thirty ways of using AI, and no one could say where the data went

The lender serves about 4 million retail and small-business borrowers through 600 branches and a digital channel. Over twelve months, teams in collections, customer service, credit, engineering, risk and HR had each found their own uses for language models: drafting credit notes, summarising call transcripts, answering policy questions, writing code.

A security review found 14 different external AI services in use, several paid on corporate cards, and at least two cases where loan application text containing names, PAN and phone numbers had been pasted into a public chatbot. Nobody could produce a list of which prompts had left the company, or what they cost.

The CIO did not want to stop the work: some of it was saving real time. The ask was for one controlled way in, where the safe choice is also the easy one for teams, customer data stays on company hardware, and the board can see usage, cost and risk on a single page.

What could not be compromised

  • Customer personal data (KYC, loan, repayment, contact details) must never be sent to a model outside the company’s data centre.
  • DPDP Act obligations: purpose limitation, data minimisation, access control and the ability to show what personal data was processed and why.
  • RBI expectations on IT governance and outsourcing: external AI providers are treated as material vendors, with contracts, audits and exit plans.
  • Teams must not need to rewrite applications every time a model changes. One API, one set of credentials per team.
  • Spend on external APIs is capped per team and visible to finance every month.
Options weighed

Four ways to give teams access to models

Each option was scored on data control, cost at 250,000 requests a day, speed for teams, and the effort to run it with a platform team of six.

OptionWhat worksWhat does notVerdict
Ban external AI, build nothingNo data leaves. Nothing to run.Teams carry on with personal accounts anyway, now hidden. Lost productivity.Rejected
One external provider, enterprise contractQuick to start. Strong models.Customer data would still leave the company. No routing by sensitivity, single-vendor lock-in.Rejected
Everything self-hostedComplete data control.Matching the best external models for low-risk tasks needs far more GPUs. Slower access to new models.Rejected
Gateway with routing by data class: self-hosted for sensitive, approved APIs for the restData control where it matters, best model for the rest, one place for cost and audit.Classification has to be reliable, and the gateway becomes critical infrastructure.Chosen
Target architecture

Every request takes the same path, and the data class decides where it can go

Applications call one endpoint with a team key. The gateway checks the team’s quota, classifies the request, removes personal data where needed, and only then picks a model. Confidential and Restricted data can only reach models on the company’s own GPUs. The network enforces the same rule: only the gateway has a route to external AI providers.

Scroll sideways to see the whole diagram →
30+ TEAMS AND APPSLLM GATEWAY, IN THE COMPANY DATA CENTREON-PREM GPU CLUSTERAPPROVED EXTERNAL APISLoan operationscredit notes, collectionsCustomer servicechat and email assistEngineering teamscode and docs assistRisk and analyticspolicy search, reportsGateway APISSO, team keys, one endpointClassify and redact PIIAadhaar, PAN, account, phonePolicy routerby data class, cost, latencyPrompt and response logredacted, access-controlledBudgets and rate limitsper team, per use caseResponse cacheexact and semantic matchModel catalogueapproved models, data classesEvaluation and reportsquality, cost, usage by teamLarge open model70B-class, 8 x H100Small open model8B-class, 2 x L40SEmbedding modelsfor RAG search, 2 x L40SExternal LLM APIszero-retention contracts2every request3data class tag4Confidential, Restricted5Public, Internal: redacted6logged1User or API trafficControl / API callData / replicationLogging / management
Numbered flows: (1) a team application calls the gateway with its own key, (2) every request is classified and personal data is detected and masked, (3) the request carries a data-class tag to the router, (4) Confidential and Restricted requests go only to self-hosted models, (5) Public and Internal requests may go to an approved external API, after redaction, (6) every prompt and response is logged in redacted form.
Building blockWhy it is there
1 Gateway APIOne OpenAI-compatible endpoint for every team. Staff sign in through company SSO; applications use a team key tied to a named use case. Teams switch models by changing a name in the request, not their code.
2 Classify and redact PIIA small classifier tags each request as Public, Internal, Confidential or Restricted, using the declared use case plus what it finds in the text. Pattern rules and an NER model detect Aadhaar, PAN, account and loan numbers, phone numbers, emails and names, and replace them with tokens before anything is logged or routed. If detection and declaration disagree, the stricter class wins.
3 Policy routerChooses the model from the catalogue by data class first, then by cost and latency. Restricted and Confidential traffic has no route to external providers at all. If the self-hosted pool is busy, requests queue; they never spill outside.
4 Self-hosted modelsA 70B-class open model on 8 x H100 for harder reasoning and long documents, an 8B-class model on L40S cards for summaries and classification, and embedding models for teams building RAG search. Served with continuous batching.
5 Approved external APIsTwo contracted providers with zero data retention, no training on prompts, and audit rights. Reached only from the gateway’s egress proxy; the firewall blocks AI provider domains from every other subnet.
6 Budgets, rate limits and cacheEach team gets a monthly token budget and requests-per-minute limit per use case. Alerts at 80%, hard stop at 100% unless the budget owner raises it. Exact and semantic caching answer repeated policy and FAQ questions without calling a model.
7 Prompt and response logEvery request is stored after redaction with team, use case, model, tokens, cost and latency. Raw text is kept only for Restricted use cases that need it for audit, encrypted and readable by two named roles.
8 Model catalogue and evaluationA register of approved models, what data classes each may receive, owners and review dates. New models and prompt changes are scored against each team’s test set before they go live.
Sizing, worked out

GPU capacity sized from request mix, not from model size alone

Request volumes come from three months of logs on the pilot teams, scaled to all 30 teams. About 70% of requests carry Confidential or Restricted data, so most traffic has to stay on company GPUs.

ItemFigureBasis
Requests~250,000 a day, peak ~15 a second30 teams, 10-hour working day, peak hour at 2x average
Tokens per request~2,500 in, ~350 outMeasured in pilot; RAG and document tasks pull the average up
Self-hosted share~70%, ~10.5 requests a second at peakConfidential and Restricted share from pilot classification
Large model8 x H100 (2 replicas x 4 GPUs)~40% of self-hosted traffic = ~4.2 req/s x 350 = ~1,500 output tokens/s at FP8; one replica handles it, two for failover
Small model and embeddings4 x L40S~6.3 req/s on the 8B model (2 GPUs) plus embeddings and the PII model (2 GPUs)
External API spend₹8 to 12 lakh a month~30% of requests, ~4.7 billion tokens a month at blended contract rates, split into team caps
Log storage~1.2 TB a year~250,000 requests x ~13 KB compressed x 365 days, kept per retention policy

Before the H100 order, the chosen 70B model is benchmarked on rented capacity with replayed pilot traffic. The cache is expected to absorb 15 to 20% of requests once policy and FAQ use cases are live, which is held as headroom rather than used to cut GPUs.

How it is delivered

From two pilot teams to thirty, with the side doors closed last

The gateway has to be the easiest way in before the other ways are closed. Teams move in waves, each with a named use case and a test set.

1

Policy and design

Weeks 1 to 4

Data classes agreed with risk, compliance and the DPO. Use-case register started. External providers assessed as material vendors.

Gate: Classification policy and provider contracts signed off.

2

Build the gateway

Weeks 5 to 10

Gateway, redaction, router, logging and budgets built. GPU servers installed and models benchmarked.

Gate: Redaction recall above 99% on a labelled set of 5,000 real prompts.

3

Pilot teams

Weeks 11 to 14

Customer service and engineering move first. Shadow classification compares declared and detected data classes.

Gate: No Restricted request routed externally in four weeks of logs.

4

Onboard all teams

Weeks 15 to 18

Remaining teams join in three waves, each with budgets, keys and an evaluation set.

Gate: Every active AI use case listed in the register with an owner.

5

Close side doors

Weeks 19 to 20

Firewall blocks AI provider domains except from the gateway. Corporate-card AI subscriptions cancelled.

Gate: Network logs show no direct AI traffic for two weeks.

Way back: Any model can be pulled from the catalogue in one change and its traffic falls back to the next approved model. If the gateway itself fails, teams lose AI features but nothing else; applications are built to degrade without them, and personal-data traffic never fails over to an external route.
Risks, handled up front

The risks that matter with a shared AI gateway

RiskWhat could happenHow the design handles it
Missed personal dataRedaction misses a name or account number in free textStricter class wins on disagreement, so a missed item still goes to a self-hosted model. Weekly sampled review of logs; recall measured against a growing labelled set.
Gateway as single point of failureAn outage takes AI out of 30 teams at onceThree gateway instances across two racks, stateless behind a load balancer. Applications must handle AI being unavailable.
Cost runawayA looping script burns a month’s budget overnightPer-minute rate limits, hard monthly caps and alerts at 80%. External spend is reported by team to finance every month.
Prompt injection through documentsA customer document tries to change a model’s instructionsDocument text is passed as data, system prompts are fixed per use case, and the gateway gives models no tools by default.
Vendor change or exitAn external provider changes terms or retentionTwo approved providers behind one API. Contracts include exit clauses; the catalogue lets traffic move in a day.
What was optimised

What was optimised

70%

Kept on company GPUs

All Confidential and Restricted traffic runs on self-hosted models. Only redacted, lower-risk work goes outside.

1

Endpoint for every team

One API and one key per use case. Swapping a model is a catalogue change, not a code change.

15 to 20%

Answered from cache

Repeated policy and FAQ questions are served without calling a model.

8B first

Right model for the job

Summaries and classification go to the small model. The 70B model is kept for work that needs it.

80%

Budget alerts before the cap

Team owners see spend daily and get a warning well before a hard stop.

0

Direct routes to AI providers

The firewall makes the gateway the only way out, so policy cannot be bypassed.

Outcomes

What the design is built to deliver

MeasureBeforeDesign target
External AI services in use14, unmanaged2 approved providers, behind the gateway
Customer data sent to external modelsUnknownNone, enforced by routing and network
AI spend visibilityScattered card paymentsMonthly report by team and use case
Time to give a new team accessWeeks of ad hoc approval2 to 3 days with a registered use case
Cost per 1,000 requestsNot measured25 to 40% lower than pilot-era external spend for the same work
Audit evidence for DPDP and RBI reviewsManual interviewsUse-case register, data-class logs and provider contracts in one pack

Targets are confirmed during the pilot on real traffic. GPU sizing and external spend depend on the final model choices and on how many use cases turn out to need the large model.

Skills this draws on

What a team needs to deliver this

LLM platform design

Gateways, routing, catalogues and APIs that let many teams use models safely through one door.

Data classification and PII redaction

Detection of Indian identifiers, NER tuning and recall testing on real prompts.

GPU serving and capacity planning

Model selection, quantisation, batching and sizing from measured request mix.

AI governance

Use-case registers, model approval and evaluation that satisfy risk, compliance and the DPO.

Regulatory mapping

Turning DPDP Act and RBI outsourcing expectations into controls an auditor can test.

Network and egress control

Proxy and firewall design so that policy is enforced by the network, not by goodwill.

About this page. This is a reference deployment: a worked design built from requirements we see repeatedly in this kind of organisation. It is not a description of a specific client. Figures are design targets and planning estimates; real numbers depend on your workloads and are confirmed during assessment. We are glad to walk through how it would apply to your environment.

Do you know how many AI services your teams are already using?

Share a rough list of AI use cases and the data they touch. We will come back with a plain view of what can safely go to external models, what must stay on your own hardware, and what a gateway would need to look like.