Retrieval metrics tell you whether the right evidence reached the model. They say nothing about what the model did with it. A model given perfect context can still add a detail that is not there, answer a neighbouring question, or answer only half of what was asked. To catch these failures we evaluate the generation step too, and the system end to end.
Three things to compare
A RAG interaction has three texts that can be compared pairwise: the question, the retrieved context and the answer. A fourth, the expected (ground-truth) answer, is available when you have labelled data. Each core metric compares a different pair.
Faithfulness (answer ↔ context)
Is every claim in the answer supported by the retrieved context? The context says the company was founded in 2003, and the answer says 2003: faithful. The answer says 2005: unfaithful, which means a hallucination despite having the right evidence. Faithfulness is the key metric for trust, because an unfaithful RAG system defeats its own purpose. Typically an LLM judge splits the answer into individual claims, checks each against the context, and reports the fraction supported.
Answer relevancy (answer ↔ question)
Does the answer address what was asked? Asked "when was Tesla founded?", an answer listing the Model S, Model 3 and Model X may be perfectly faithful to the context and still be irrelevant. Relevancy penalises padding, evasion and answering the wrong question.
Contextual precision, recall and relevancy (context ↔ question / expected answer)
These are the LLM-judged versions of the retrieval metrics from the previous chapter, and they need no chunk-ID labels:
- Contextual recall: does the retrieved context contain everything needed to produce the expected answer? Missing facts mean low recall.
- Contextual precision: are the relevant pieces of context ranked above the irrelevant ones?
- Contextual relevancy: what fraction of the retrieved context is relevant to the question at all?
Answer correctness (answer ↔ expected answer)
Does the answer match the ground truth? This is the bottom line from the user's point of view. It needs a labelled expected answer, and it is judged on meaning, not on exact wording, since two correct answers rarely share their phrasing.
Reading the metrics together
The value of having several metrics is diagnosis. Each failure pattern points to a different component:
| Pattern | Diagnosis | Fix in |
|---|---|---|
| Low contextual recall | Retrieval missed the evidence | Chunking, hybrid search, query transformation |
| Good recall, low contextual precision | Evidence found but ranked low | Reranking |
| Good context, low faithfulness | Model ignores or embellishes the evidence | Prompt ("answer only from context"), model choice, temperature |
| Faithful but low answer relevancy | Model answers a neighbouring question | Prompt, query understanding, decomposition |
| Everything high, low correctness | Labels, or the knowledge base itself, may be wrong or outdated | Data (Module 6) |
LLM-as-judge
Faithfulness and relevancy cannot be computed by string matching, because they are judgements of meaning. Evaluation frameworks such as DeepEval and RAGAS use an LLM as the judge: for each test case they prompt a capable model with a rubric ("extract the claims; for each, is it supported by this context?") and turn its verdicts into a score, usually with a written reason you can read.
Practical points:
- A test case bundles the question, the system's actual answer, the expected answer and the retrieved context. The context is needed because several metrics judge it directly.
- Each metric has a threshold, for example 0.7, which is the pass mark for that metric's 0–1 score. It is not a required "70% word overlap" with the expected answer. A case passes the metric if its score is at or above the threshold.
- The judge is itself a model, so it is noisy and biased: it may favour longer answers or its own phrasing. Use a strong judge model, keep the rubric fixed, and spot-check a sample of verdicts by hand, especially when a score moves unexpectedly.
- Judging costs tokens: (number of cases) × (number of metrics) × (several LLM calls each). Run the full suite before releases, and a smaller smoke test on every change.
Building the golden dataset
Every metric needs test questions, and the correctness metrics need expected answers. Writing hundreds by hand is slow, so frameworks provide synthesizers: give them your documents and an LLM generates (question, expected answer, source context) triples, a starter golden dataset in minutes.
Synthetic goldens are a starting point, not the finish line:
- Review them. Delete trivial or ambiguous questions and fix wrong expected answers. A judge comparing against a wrong label produces confident nonsense.
- Add real questions. Production logs show how users actually phrase things: vague, compound, misspelt. Those are the questions that break systems.
- Cover the hard cases deliberately: questions whose answers span several chunks, questions about codes and IDs, questions that should be answered "I don't know" because the answer isn't in the corpus, and multi-part questions.
- Version it alongside your code, so scores remain comparable across experiments.
Treat the golden set like a test set in machine learning. Never tune prompts on the exact questions you report scores for, or the scores will flatter you. Keep a held-out slice you look at only to confirm improvements.