By now you have a set of options: chunk sizes, embedding models, hybrid weights, filters, rerankers and query rewriting. Each one might help. Without measurement, changing them is guesswork. You try a question you remember failing, it now works, and you ship a change that quietly broke ten other questions.
A RAG system has two stages to evaluate: retrieval (did we fetch the right evidence?) and generation (did the model use it well?). Start with retrieval, for a simple reason: if the right chunk never reaches the model, the answer cannot be right, however good the model is. Retrieval metrics are also cheap, deterministic and fast to compute. This chapter covers them, and the next covers generation.
The setup: a labelled evaluation set
Every metric below needs the same input: a set of test questions, each paired with the IDs of the chunks that are relevant to it (and, for graded metrics, how relevant). For example:
| Question | Relevant chunk IDs |
|---|---|
| What is the notice period? | hr_notice_c1, hr_notice_c2 |
| Can I carry forward leave? | hr_leave_c4 |
You build this once, by hand or semi-automatically (next chapter), and reuse it for every experiment. For each question you run your retriever, take the ranked list of the top-k results, and compare it with the labels.
Throughout, k is how many results you keep: the same number you would pass to the model or the reranker.
Recall@k: did we find what exists?
If a question has 5 relevant chunks and the top 5 results contain 3 of them, recall@5 = 3/5 = 60%. Recall answers the question "of everything that should have been found, how much was?" For RAG it is usually the most important retrieval metric, because a missing chunk is a missing fact.
Precision@k: how much of what we returned is useful?
If the top 5 contain 2 relevant chunks, precision@5 = 2/5 = 40%. Precision measures noise: low precision means the prompt is padded with irrelevant text that costs tokens and can distract the model.
The two pull against each other. Raising k usually increases recall (more chances to catch the relevant chunks) and lowers precision (more filler). The difference between them is a favourite interview question: recall's denominator is the number of relevant items; precision's denominator is k.
Hit rate@k: did we get at least one?
Hit rate@k is 1 for a question if at least one relevant chunk appears in the top k, and 0 otherwise, averaged over all questions. If 4 of 5 test questions have a relevant chunk somewhere in their top k, hit rate = 80%. It ignores position and quantity. It is a blunt but intuitive health check ("how often does the model see any useful evidence?"), and it suits questions with exactly one relevant chunk.
MRR: how high is the first relevant result?
Position matters: a relevant chunk at rank 1 is more useful than one at rank 9 (and after reranking only the top few survive). The reciprocal rank of a question is of its first relevant result, or 0 if there is none. Mean reciprocal rank (MRR) averages that over questions.
Worked example. Three questions whose first relevant chunk appears at ranks 1, 4 and 2:
MRR is ideal when one good chunk is enough to answer the question (FAQ-style lookups).
MAP: ranking quality when several chunks are relevant
Average precision (AP) for one question looks at every position where a relevant chunk appears, takes the precision at that point, and averages those precisions over all relevant chunks (a relevant chunk never retrieved contributes 0).
Worked example. The top 5 are [✓, ✗, ✗, ✓, ✗] and there are 2 relevant chunks in total.
- At rank 1, the first relevant chunk: precision@1 = 1/1 = 1.0.
- At rank 4, the second relevant chunk: precision@4 = 2/4 = 0.5.
- AP = (1.0 + 0.5) / 2 = 0.75.
Now suppose a question has 3 relevant chunks and its top 5 are [✗, ✓, ✓, ✗, ✗]: AP = (1/2 + 2/3 + 0) / 3 ≈ 0.389. The third relevant chunk was never found, and that pulls the score down. Mean average precision (MAP) is the mean of AP across questions. It rewards placing all relevant chunks early.
NDCG: graded relevance with a position discount
Some chunks are perfect, some partly useful, some irrelevant. NDCG (normalised discounted cumulative gain) uses graded relevance labels, such as 3 = highly relevant, 2 = relevant, 1 = marginal, 0 = irrelevant, and discounts gains logarithmically with rank:
where IDCG is the DCG of the ideal ordering (the same chunks sorted best-first). The discount is 1 at rank 1, 1.585 at rank 2, 2 at rank 3, and so on. Note the : dividing by at rank 1 would be undefined, which is a common slip when the formula is written from memory.
Worked example. Retrieved relevance in order: [3, 2, 0, 1, 3].
The ideal ordering is [3, 3, 2, 1, 0]:
NDCG = 5.853 / 6.323 ≈ 0.926. Most of the value is there, but a highly relevant chunk sitting at rank 5 cost about 7%. NDCG reaches 1.0 only when the ordering is ideal.
Which metric when?
| Metric | Rewards | Use when |
|---|---|---|
| Recall@k | Finding all relevant chunks | Almost always: missing evidence means a missing fact |
| Precision@k | Low noise | Context budget is tight, or noise confuses the model |
| Hit rate@k | Getting at least one | Single-answer lookups; quick health check |
| MRR | First relevant result early | One good chunk suffices (FAQ, support) |
| MAP | All relevant results early (binary labels) | Multi-chunk answers |
| NDCG | Best results earliest (graded labels) | Mixed-quality candidates; evaluating rerankers |
A practical dashboard tracks recall@k (are we finding it?), MRR or NDCG (are we ranking it well?) and precision@k (how much noise do we send?), for every configuration you try.
Using the numbers
Metrics turn arguments into experiments. Low recall at large k points upstream: chunking, the embedding model, hybrid search or query transformation. High recall@50 but low recall@3 points to ranking, so add or tune a reranker. Low precision with good recall suggests lowering k or reranking harder. Re-run the whole evaluation set after every change, and keep the per-question results so you can see which questions improved and which regressed.
Tap results to grade them and see recall, precision, MRR, MAP and NDCG update, with the arithmetic shown.