Measuring Retrieval: Recall, Precision, MRR, MAP and NDCG

If retrieval fails, nothing downstream can save the answer, so measure it first. This chapter defines the standard retrieval metrics, works through each by hand, explains what each one rewards, and shows how to build the labelled dataset they need.

Advanced RAG Masterclass

By now you have a set of options: chunk sizes, embedding models, hybrid weights, filters, rerankers and query rewriting. Each one might help. Without measurement, changing them is guesswork. You try a question you remember failing, it now works, and you ship a change that quietly broke ten other questions.

A RAG system has two stages to evaluate: retrieval (did we fetch the right evidence?) and generation (did the model use it well?). Start with retrieval, for a simple reason: if the right chunk never reaches the model, the answer cannot be right, however good the model is. Retrieval metrics are also cheap, deterministic and fast to compute. This chapter covers them, and the next covers generation.

The setup: a labelled evaluation set

Every metric below needs the same input: a set of test questions, each paired with the IDs of the chunks that are relevant to it (and, for graded metrics, how relevant). For example:

QuestionRelevant chunk IDs
What is the notice period?hr_notice_c1, hr_notice_c2
Can I carry forward leave?hr_leave_c4

You build this once, by hand or semi-automatically (next chapter), and reuse it for every experiment. For each question you run your retriever, take the ranked list of the top-k results, and compare it with the labels.

Throughout, k is how many results you keep: the same number you would pass to the model or the reranker.

Recall@k: did we find what exists?

Recall@k=relevant chunks in the top kall relevant chunks for the question.\text{Recall@}k = \frac{\text{relevant chunks in the top } k}{\text{all relevant chunks for the question}}.

If a question has 5 relevant chunks and the top 5 results contain 3 of them, recall@5 = 3/5 = 60%. Recall answers the question "of everything that should have been found, how much was?" For RAG it is usually the most important retrieval metric, because a missing chunk is a missing fact.

Precision@k: how much of what we returned is useful?

Precision@k=relevant chunks in the top kk.\text{Precision@}k = \frac{\text{relevant chunks in the top } k}{k}.

If the top 5 contain 2 relevant chunks, precision@5 = 2/5 = 40%. Precision measures noise: low precision means the prompt is padded with irrelevant text that costs tokens and can distract the model.

The two pull against each other. Raising k usually increases recall (more chances to catch the relevant chunks) and lowers precision (more filler). The difference between them is a favourite interview question: recall's denominator is the number of relevant items; precision's denominator is k.

Hit rate@k: did we get at least one?

Hit rate@k is 1 for a question if at least one relevant chunk appears in the top k, and 0 otherwise, averaged over all questions. If 4 of 5 test questions have a relevant chunk somewhere in their top k, hit rate = 80%. It ignores position and quantity. It is a blunt but intuitive health check ("how often does the model see any useful evidence?"), and it suits questions with exactly one relevant chunk.

MRR: how high is the first relevant result?

Position matters: a relevant chunk at rank 1 is more useful than one at rank 9 (and after reranking only the top few survive). The reciprocal rank of a question is 1/rank1/\text{rank} of its first relevant result, or 0 if there is none. Mean reciprocal rank (MRR) averages that over questions.

Worked example. Three questions whose first relevant chunk appears at ranks 1, 4 and 2:

MRR=13(11+14+12)=1.753≈0.583.\text{MRR} = \frac{1}{3}\left(\frac11 + \frac14 + \frac12\right) = \frac{1.75}{3} \approx 0.583.

MRR is ideal when one good chunk is enough to answer the question (FAQ-style lookups).

MAP: ranking quality when several chunks are relevant

Average precision (AP) for one question looks at every position where a relevant chunk appears, takes the precision at that point, and averages those precisions over all relevant chunks (a relevant chunk never retrieved contributes 0).

Worked example. The top 5 are [✓, ✗, ✗, ✓, ✗] and there are 2 relevant chunks in total.

  • At rank 1, the first relevant chunk: precision@1 = 1/1 = 1.0.
  • At rank 4, the second relevant chunk: precision@4 = 2/4 = 0.5.
  • AP = (1.0 + 0.5) / 2 = 0.75.

Now suppose a question has 3 relevant chunks and its top 5 are [✗, ✓, ✓, ✗, ✗]: AP = (1/2 + 2/3 + 0) / 3 ≈ 0.389. The third relevant chunk was never found, and that pulls the score down. Mean average precision (MAP) is the mean of AP across questions. It rewards placing all relevant chunks early.

NDCG: graded relevance with a position discount

Some chunks are perfect, some partly useful, some irrelevant. NDCG (normalised discounted cumulative gain) uses graded relevance labels, such as 3 = highly relevant, 2 = relevant, 1 = marginal, 0 = irrelevant, and discounts gains logarithmically with rank:

DCG@k=∑i=1krelilog⁡2(i+1),NDCG@k=DCG@kIDCG@k,\text{DCG@}k = \sum_{i=1}^{k} \frac{\text{rel}_i}{\log_2(i+1)}, \qquad \text{NDCG@}k = \frac{\text{DCG@}k}{\text{IDCG@}k},

where IDCG is the DCG of the ideal ordering (the same chunks sorted best-first). The log⁡2(i+1)\log_2(i+1) discount is 1 at rank 1, 1.585 at rank 2, 2 at rank 3, and so on. Note the +1+1: dividing by log⁡(1)=0\log(1)=0 at rank 1 would be undefined, which is a common slip when the formula is written from memory.

Worked example. Retrieved relevance in order: [3, 2, 0, 1, 3].

DCG=31+21.585+02+12.322+32.585=3+1.262+0+0.431+1.161=5.853.\text{DCG} = \tfrac{3}{1} + \tfrac{2}{1.585} + \tfrac{0}{2} + \tfrac{1}{2.322} + \tfrac{3}{2.585} = 3 + 1.262 + 0 + 0.431 + 1.161 = 5.853.

The ideal ordering is [3, 3, 2, 1, 0]:

IDCG=3+31.585+22+12.322+0=3+1.893+1+0.431=6.323.\text{IDCG} = 3 + \tfrac{3}{1.585} + \tfrac{2}{2} + \tfrac{1}{2.322} + 0 = 3 + 1.893 + 1 + 0.431 = 6.323.

NDCG = 5.853 / 6.323 ≈ 0.926. Most of the value is there, but a highly relevant chunk sitting at rank 5 cost about 7%. NDCG reaches 1.0 only when the ordering is ideal.

A ranked list of five results marked relevant or not, annotated with how recall@5, precision@5, hit rate, reciprocal rank, average precision and NDCG each read it
One ranked list, six readings. Recall and precision count; hit rate checks for any; MRR looks at the first hit; MAP and NDCG reward putting all the good results early.

Which metric when?

MetricRewardsUse when
Recall@kFinding all relevant chunksAlmost always: missing evidence means a missing fact
Precision@kLow noiseContext budget is tight, or noise confuses the model
Hit rate@kGetting at least oneSingle-answer lookups; quick health check
MRRFirst relevant result earlyOne good chunk suffices (FAQ, support)
MAPAll relevant results early (binary labels)Multi-chunk answers
NDCGBest results earliest (graded labels)Mixed-quality candidates; evaluating rerankers

A practical dashboard tracks recall@k (are we finding it?), MRR or NDCG (are we ranking it well?) and precision@k (how much noise do we send?), for every configuration you try.

Using the numbers

Metrics turn arguments into experiments. Low recall at large k points upstream: chunking, the embedding model, hybrid search or query transformation. High recall@50 but low recall@3 points to ranking, so add or tune a reranker. Low precision with good recall suggests lowering k or reranking harder. Re-run the whole evaluation set after every change, and keep the per-question results so you can see which questions improved and which regressed.

Try it yourself
RAG Lab: retrieval metrics →

Tap results to grade them and see recall, precision, MRR, MAP and NDCG update, with the arithmetic shown.

MediumEvaluationMetrics

A question has 4 relevant chunks; the top 5 results contain 2 of them, at ranks 2 and 5. Compute recall@5, precision@5, reciprocal rank and AP.

EasyEvaluationInterview

What is the difference between recall@k and precision@k?

HardEvaluationNDCG

Why does NDCG divide by IDCG, and why the log₂(i+1) discount?