Home / RAG / Measuring answers
RAG guide

Measuring answers

Without measurement, every change to a RAG system is a guess. A small, well-built test set turns quality into a number you can track and improve.

Test set100-200 real questions
Check 1Right passage found?
Check 2Answer faithful?
RunOn every change

What to measure

MeasureQuestion it answersHow
Retrieval hit rateDid the right passage come back at all?Automatic, using the source marked in the test set
FaithfulnessDoes the answer stick to the passages, without inventing?Automatic scoring plus expert review of a sample
CorrectnessIs the answer actually right?Compared with the agreed answer
RefusalsDoes it say "I don't know" when the answer is not in the documents?Include questions with no answer in the test set
User feedbackDo people find it useful?Thumbs up or down, plus comments

Building the test set

  • Collect real questions from help desks, email and chat, not invented ones.
  • For each, record the right answer and the document it comes from.
  • Include hard questions, questions with no answer, and questions that need permissions.
  • Re-run it on every change to documents, prompts, models or search settings. Tools such as Ragas help automate scoring.

Want answers from your own documents?

Tell us where the documents live, who should be able to ask, and ten questions people ask today. We will come back with a plain view of what it takes to answer them well, and safely.