Home / Deployment scenarios / AI copilot for a service contact centre
Reference deployment Contact centre AI

An AI copilot for 600 service agents, in Hindi, English and three regional languages

A consumer-durables brand runs a 600-seat contact centre for repairs, installations and warranty questions across more than 1,400 product models. Agents search PDFs while customers wait, type notes after every call, and quality analysts hear about 2% of conversations. This design adds a copilot that listens alongside the agent, finds the answer in the manuals, writes the call notes and scores every call, with all of it running on the company’s own GPUs.

SectorConsumer durables, after-sales service
Scale600 seats, ~30,000 calls/day
Programme22 weeks
ModelSelf-hosted speech and LLMs
The situation

Long calls, thin notes and a quality team that hears one call in fifty

The service contact centre handles about 30,000 calls a day: booking repairs, explaining error codes, checking warranty and arranging installations for air conditioners, refrigerators, washing machines and televisions. Calls come in Hindi, English and code-mixed Hinglish, with a growing share in Tamil, Telugu and Bengali.

Average handle time is about 6.5 minutes, of which around a minute is the agent searching service manuals and bulletins, and another 70 seconds is typing notes after the call. Notes are short and inconsistent, so when a customer calls back, the next agent often starts from zero. About 18% of repair visits are repeat visits, many traced back to wrong first diagnosis on the phone.

Twenty-four quality analysts listen to about 600 calls a day, roughly 2%. Coaching depends on which calls happen to be picked. The head of customer service wanted every call heard, agents helped in the moment, and no customer card or phone number sent to an outside AI service.

What could not be compromised

  • Call audio and transcripts must stay on company-controlled infrastructure. No external speech or LLM APIs.
  • Card numbers must never be stored in transcripts or recordings, in line with PCI DSS. Phone, Aadhaar and address details must be masked in everything the AI stores.
  • Live suggestions must reach the agent within about 2 seconds of the customer finishing a sentence, or they are ignored.
  • Must work in Hindi, English and code-mixed speech from day one, with Tamil, Telugu and Bengali added in phases.
  • The existing telephony platform, agent desktop and CRM stay. The copilot integrates with them; it does not replace them.
Options weighed

Three ways to add AI to the contact centre

Options were compared on 400 recorded calls across five languages, scored on transcript accuracy, answer quality, data control and five-year cost at 600 seats.

OptionWhat worksWhat does notVerdict
Contact-centre vendor’s AI add-onQuick to switch on inside the existing platform.Audio processed in the vendor’s cloud. Weak on regional languages and code-mixed speech. Per-seat fees at 600 seats add up.Rejected
Cloud speech and LLM APIs, built in-houseStrong models, no GPUs to run.Card and phone numbers in audio would leave the company before masking. Per-minute speech costs at 30,000 calls a day.Rejected
Self-hosted speech and LLMs, integrated with existing systemsAll audio stays inside. Models tuned on the company’s own calls and product vocabulary. Cost fixed by hardware.GPU capacity and model tuning to manage. Regional languages need labelled audio.Chosen
Target architecture

One audio stream, two paths: live help during the call, notes and quality after it

The telephony platform forks the audio of every call to the copilot. Speech-to-text runs on both channels as the call happens, and personal data is masked before anything else sees the text. During the call, the assist service finds answers in the service manuals and shows them to the agent. After hang-up, the summariser writes notes into the CRM and every call is scored against the quality form.

Scroll sideways to see the whole diagram →
CUSTOMERS AND TELEPHONYAGENTSAI COPILOT, SELF-HOSTEDBUSINESS SYSTEMSQUALITY TEAMKNOWLEDGE SOURCESCustomersphone, 5 languagesTelephony and ACDIVR, routing, SIPREC forkAgent desktop600 seats, assist panelTeam leaderslive escalation alertsStreaming speech-to-text12 x L40S, both channelsPII maskingcard, phone, Aadhaar numbersLive agent assistanswers with manual citationsSelf-hosted LLMs8B assist, 32B QA, 8 x L40STranscript storemasked text, per-call indexCall summarisernotes within 30 s of hang-upQuality scoring100% of calls, 22-point formManual search indexchunked, Hindi and EnglishCall recordingscard digits mutedCRMcases, summaries, tagsQA and supervisors24 analysts, coachingService manuals1,400 models, bulletinscalls2live transcriptmasked textescalation riskpromptson hang-up5every call1346User or API trafficData / replicationControl / API callException / alertLogging / managementScheduled copy
Numbered flows: (1) the telephony platform forks audio from both channels of every call, (2) streaming speech-to-text produces a live transcript, which is masked before use, (3) agent assist retrieves answers from the manual index and shows them on the agent desktop, (4) on hang-up, a summary and disposition go into the CRM case, (5) every call is scored against the quality form, (6) low scores and compliance flags go to QA analysts and supervisors.
Building blockWhy it is there
1 Audio fork from telephonyThe existing platform sends a copy of each call’s audio over SIPREC, with agent and customer on separate channels. The live call path is untouched, so if the copilot fails, calls carry on as today.
2 Streaming speech-to-textOpen speech models fine-tuned on 600 hours of the company’s own labelled calls, including product names, model numbers and error codes. Handles Hindi, English and code-mixed speech, with Tamil, Telugu and Bengali added as their training data is ready.
3 PII maskingPattern rules and a small entity model replace card, phone, Aadhaar and address details with tokens in the transcript within a second. Card payments move to IVR keypad entry with recording paused; the same digits are also muted in stored audio as a backstop.
4 Live agent assistDetects the customer’s issue as they describe it, searches the manual index and shows the agent two or three short answers with the manual page they came from. Agents rate each suggestion with one click, which feeds evaluation.
5 Manual search indexService manuals, error-code tables and technical bulletins for 1,400 models, split into sections, indexed with multilingual embeddings and refreshed nightly from the document system.
6 Call summariserWrites a structured note into the CRM case within 30 seconds of hang-up: issue, model, steps tried, outcome, next action. The agent reviews and saves it instead of typing from scratch.
7 Quality scoringA larger model scores every call against the company’s 22-point quality form, with the transcript lines that support each score. Analysts review the low scores, compliance flags and a random sample, rather than random calls only.
8 Self-hosted LLMsAn 8B-class multilingual model for assist and summaries, where speed matters, and a 32B-class model for quality scoring, where depth matters. All run on L40S GPUs in the company’s data centre.
Sizing, worked out

GPUs sized from concurrent calls at peak, not from seat count

Peak load is the busiest hour on a Monday after a long weekend, measured from twelve months of ACD reports. Throughput per GPU comes from a benchmark on rented hardware with the company’s own recorded calls.

ItemFigureBasis
Peak concurrent calls~510600 seats x 85% occupancy in the peak hour
Live audio streams~1,020Agent and customer channels transcribed separately
Speech-to-text GPUs12 x L40S (3 servers)~130 streams per GPU in benchmark, so 8 GPUs at peak; 12 keeps peak covered with one server down
Assist requests~11 a second at peak510 calls x one query every ~45 seconds of conversation
Assist and summary model4 x L40S8B-class at FP8, ~2,000 tokens in and ~150 out per query; ~22,000 input tokens/s at peak needs 3 GPUs, plus 1 spare. Summaries add ~1 a second
Quality scoring model4 x L40S (2 replicas x 2 GPUs)30,000 calls x ~500 output tokens = ~15 million tokens a day at ~1,600 tokens/s, about 3 hours of GPU time spread through the day
Storage~0.4 TB a year~30,000 masked transcripts, summaries and scores a day at ~35 KB each

Total: 20 x L40S across five servers, estimated at ₹3 to 4.5 crore including networking and three years of support. Speech-to-text is the largest load and scales with concurrent calls, so adding seats is a matter of adding speech servers.

How it is delivered

Post-call first, then live help, then more languages

Summaries and quality scoring carry little risk because a person reviews them. Live assist goes to a pilot floor only once speech accuracy is proven on real calls.

1

Data and benchmark

Weeks 1 to 5

600 hours of call audio transcribed and labelled. Models benchmarked on rented GPUs. Card capture moved to IVR keypad entry.

Gate: Word error rate under 15% on Hindi, English and code-mixed test sets.

2

Post-call AI

Weeks 6 to 11

Audio fork, speech-to-text, masking, summaries and quality scoring live for all 600 seats. Agents review summaries before saving.

Gate: Analysts agree with AI scores on 85% of a 1,000-call sample.

3

Live assist pilot

Weeks 12 to 16

Assist panel switched on for 60 agents on two product lines. Suggestion ratings and handle time tracked against a control group.

Gate: Handle time down at least 30 seconds with no drop in resolution.

4

All seats

Weeks 17 to 19

Assist rolled out to all teams in three waves, with team leaders trained on escalation alerts.

Gate: Suggestion usefulness rating above 70% across teams.

5

Regional languages

Weeks 20 to 22

Tamil, Telugu and Bengali speech models added from labelled data collected during earlier phases.

Gate: Word error rate under 18% on each regional test set.

Way back: The copilot sits beside the call, not in it. Assist, summaries and scoring can each be switched off per team in seconds, and agents carry on with the desktop and CRM exactly as before. Earlier model versions stay deployed for one release so a regression can be reversed in minutes.
Risks, handled up front

What could go wrong, and what is already in the plan

RiskWhat could happenHow the design handles it
Card numbers in transcriptsA customer reads out a card number despite the IVR flowMasking runs before storage or any model sees the text, and the digits are muted in recordings. Weekly scans of stored transcripts for card patterns.
Wrong advice to customersAssist suggests a step that does not apply to that modelAnswers only come from indexed manuals with the source shown. No answer is shown when retrieval confidence is low.
Poor accuracy on accents and code-mixingTranscripts are unreliable for some regionsFine-tuning on the company’s own calls, word error rate tracked by language and region, and regional languages added only once they meet the threshold.
Agents ignore or distrust the panelAssist is switched on but not usedPilot with agent feedback built in, suggestions limited to two or three lines, and usefulness ratings reviewed weekly with team leaders.
Quality scores seen as unfairAgents dispute AI scores used in reviewsEvery score links to the transcript lines behind it. Analysts confirm any score that affects an appraisal, and agents can raise a dispute.
What was optimised

What was optimised

100%

Calls quality-scored

Every call is scored, so analysts spend their time on the calls that need attention instead of random ones.

2 models

Right model for each job

A fast 8B model for live assist and summaries, a deeper 32B model for quality scoring off the live path.

130

Streams per speech GPU

Streaming models with batching and separate channels keep 1,020 live streams on 12 GPUs with headroom.

0

Audio sent outside

Speech, models, index and transcripts all run on the company’s own servers.

30 s

Notes after hang-up

Agents review and save a structured summary instead of typing notes.

1 fork

No change to the call path

The copilot listens to a copy of the audio, so the contact centre works normally if it is down.

Outcomes

What the design is built to deliver

MeasureBeforeDesign target
Average handle time~6.5 minutes45 to 75 seconds shorter
After-call work~70 seconds of typing15 to 20 seconds to review and save
Calls quality-reviewed~2%100% scored, with analyst review of flagged calls
Repeat repair visits~18%13 to 15%, from better first diagnosis on the phone
Card and phone numbers in stored transcriptsNot checkedMasked before storage, scanned weekly
New agent time to proficiency~8 weeks5 to 6 weeks with assist and call-level coaching

Targets are confirmed against a control group during the pilot. Handle-time savings depend heavily on how much of today’s call time is spent searching for information, which is measured in the first phase.

Skills this draws on

What a team needs to deliver this

Speech AI

Streaming speech-to-text for Indian languages and code-mixed speech, fine-tuned on real call audio.

Contact-centre integration

Audio forking, agent desktop panels and CRM write-back without touching the live call path.

RAG on technical documents

Indexing service manuals and bulletins so answers come with their source and confidence.

GPU capacity planning

Sizing speech and LLM serving from concurrent calls, benchmarked on your own audio.

PII and PCI controls

Masking, muting and payment-flow changes that keep card and personal data out of AI systems.

Quality and evaluation

AI scoring calibrated against analysts, with evidence lines and dispute handling.

About this page. This is a reference deployment: a worked design built from requirements we see repeatedly in this kind of organisation. It is not a description of a specific client. Figures are design targets and planning estimates; real numbers depend on your workloads and are confirmed during assessment. We are glad to walk through how it would apply to your environment.

How many of your calls does anyone actually hear?

Share your seat count, call volumes, languages and the systems your agents use. We will come back with a plain view of what a copilot could take on, what it would need in GPUs, and how to prove it on a pilot floor first.