LLM guide
Governance and LLMOps
Once an LLM is in use, someone has to answer simple questions: who can use it, what data it sees, how good its answers are this month, and what happens when a model changes. Governance is the set of controls that makes those answers easy.
MeasureQuality, monthly
ControlAccess by group
ProtectPersonal data
RecordEvery request
The controls we put in place
| Area | What it means in practice |
|---|---|
| Access | Use is tied to company sign-in. Groups decide which models and data each team may use. |
| Data protection | Personal data is masked where it is not needed. Retention of prompts and answers follows your policy and, in India, the Digital Personal Data Protection Act, 2023. |
| Quality | A test set is re-run on every model or prompt change, and a sample of live answers is reviewed every month. |
| Safety | Filters for harmful content, and defences against prompt injection, which is the top risk in the OWASP Top 10 for LLM applications. |
| Audit | Who asked what, which model answered, and what sources were used, kept for an agreed period. |
| Cost | Usage and cost per team, with quotas and alerts. |
| Change | New models and prompts go through test, approval and gradual rollout, with a quick way back. |
Measuring quality without guesswork
- Offline: the test set, scored automatically and by reviewers, before every change.
- Online: thumbs up or down from users, plus a small weekly sample reviewed by experts.
- Drift: watch for falling scores or rising complaints after model, prompt or data changes.
Typical tools
| Area | Common choices |
|---|---|
| Tracing and evaluation | Langfuse, OpenTelemetry, Ragas for retrieval quality |
| Personal data | Microsoft Presidio, custom rules for Indian identifiers |
| Guardrails | NVIDIA NeMo Guardrails, gateway policies |
| Cost and usage | Gateway metrics with Prometheus and Grafana |
More LLM guides: Choosing a model · Private hosting · Fine-tuning · LLM FAQ · Use case: claims summaries
Starting with LLMs, or stuck after a pilot?
Tell us the task you want AI to help with and any rules about where your data can go. We will come back with a plain recommendation: which kind of model, where to run it, roughly what it costs, and how to know if it is working.