Summarising claim files with a private LLM
How we would approach an insurer that wants claims handlers to spend less time reading and more time deciding, without sending health and personal data outside its own network.
This is a composite example built from common enterprise requirements. It is not a specific client engagement, and the figures are design targets for the scenario, not measured results.
Long files, short on time
Each health or motor claim arrives with 30 to 80 pages: forms, hospital bills, discharge summaries, surveyor reports and emails. Handlers spend much of their day reading before they can decide. The company wanted a one-page summary for each file, with the key facts and anything unusual flagged. Health and identity data could not leave its own infrastructure.
What we would do
- Define the summary with handlers. Ten fields every summary must contain, and the red flags they look for.
- Build a test set. 200 past claims, anonymised, with summaries written by senior handlers.
- Compare three open models at two sizes, scored by handlers on accuracy, missing facts and usefulness.
- Choose the smallest model that passes. In this scenario, a 70B-class model at 8-bit, served on the insurer's own GPUs.
- No fine-tuning at first. A clear prompt with two worked examples met the bar. Fine-tuning stays an option if format drifts.
- Guardrails. Personal identifiers masked in logs, every summary linked to its source pages, and the handler always makes the decision.
Targets and how they are checked
| Measure | Before | Target | Checked by |
|---|---|---|---|
| Reading time per claim | About 25 minutes | About 10 minutes | Time study on a sample of claims |
| Summary accuracy | Not applicable | 95% of key facts correct | Weekly review of 5% of summaries |
| Data leaving the network | Not applicable | None | Network controls and audit logs |
What we would flag
- Scanned documents. Poor scans need good OCR first, or the model summarises garbage.
- Handlers must still read the flagged pages. The summary speeds up reading; it does not replace judgement.
- The test set needs owners. Someone in claims must keep it current as products and rules change.
Starting with LLMs, or stuck after a pilot?
Tell us the task you want AI to help with and any rules about where your data can go. We will come back with a plain recommendation: which kind of model, where to run it, roughly what it costs, and how to know if it is working.