RAG guide
Measuring answers
Without measurement, every change to a RAG system is a guess. A small, well-built test set turns quality into a number you can track and improve.
Test set100-200 real questions
Check 1Right passage found?
Check 2Answer faithful?
RunOn every change
What to measure
| Measure | Question it answers | How |
|---|---|---|
| Retrieval hit rate | Did the right passage come back at all? | Automatic, using the source marked in the test set |
| Faithfulness | Does the answer stick to the passages, without inventing? | Automatic scoring plus expert review of a sample |
| Correctness | Is the answer actually right? | Compared with the agreed answer |
| Refusals | Does it say "I don't know" when the answer is not in the documents? | Include questions with no answer in the test set |
| User feedback | Do people find it useful? | Thumbs up or down, plus comments |
Building the test set
- Collect real questions from help desks, email and chat, not invented ones.
- For each, record the right answer and the document it comes from.
- Include hard questions, questions with no answer, and questions that need permissions.
- Re-run it on every change to documents, prompts, models or search settings. Tools such as Ragas help automate scoring.
More RAG guides: How RAG is built · Preparing documents · Better retrieval · Permissions · RAG FAQ · Use case: HR policy assistant
Want answers from your own documents?
Tell us where the documents live, who should be able to ask, and ten questions people ask today. We will come back with a plain view of what it takes to answer them well, and safely.