Evaluating the Whole Pipeline: Faithfulness, Relevancy and LLM Judges

Good retrieval is necessary but not sufficient: the model can still ignore the evidence, invent details or answer a different question. This chapter covers the core end-to-end metrics, how LLM-as-judge frameworks such as DeepEval and RAGAS compute them, and how to build and use a golden dataset.

Advanced RAG Masterclass

Retrieval metrics tell you whether the right evidence reached the model. They say nothing about what the model did with it. A model given perfect context can still add a detail that is not there, answer a neighbouring question, or answer only half of what was asked. To catch these failures we evaluate the generation step too, and the system end to end.

Three things to compare

A RAG interaction has three texts that can be compared pairwise: the question, the retrieved context and the answer. A fourth, the expected (ground-truth) answer, is available when you have labelled data. Each core metric compares a different pair.

Triangle with question, retrieved context and answer at the corners; edges labelled context relevancy (question–context), faithfulness (context–answer) and answer relevancy (question–answer); a fourth node, expected answer, connects to the answer by answer correctness and to the context by contextual recall
Each metric checks one relationship. Faithfulness: is the answer supported by the context? Answer relevancy: does it address the question? Context metrics: did retrieval bring the right evidence? Correctness: does it match the ground truth?

Faithfulness (answer ↔ context)

Is every claim in the answer supported by the retrieved context? The context says the company was founded in 2003, and the answer says 2003: faithful. The answer says 2005: unfaithful, which means a hallucination despite having the right evidence. Faithfulness is the key metric for trust, because an unfaithful RAG system defeats its own purpose. Typically an LLM judge splits the answer into individual claims, checks each against the context, and reports the fraction supported.

Answer relevancy (answer ↔ question)

Does the answer address what was asked? Asked "when was Tesla founded?", an answer listing the Model S, Model 3 and Model X may be perfectly faithful to the context and still be irrelevant. Relevancy penalises padding, evasion and answering the wrong question.

Contextual precision, recall and relevancy (context ↔ question / expected answer)

These are the LLM-judged versions of the retrieval metrics from the previous chapter, and they need no chunk-ID labels:

  • Contextual recall: does the retrieved context contain everything needed to produce the expected answer? Missing facts mean low recall.
  • Contextual precision: are the relevant pieces of context ranked above the irrelevant ones?
  • Contextual relevancy: what fraction of the retrieved context is relevant to the question at all?

Answer correctness (answer ↔ expected answer)

Does the answer match the ground truth? This is the bottom line from the user's point of view. It needs a labelled expected answer, and it is judged on meaning, not on exact wording, since two correct answers rarely share their phrasing.

Reading the metrics together

The value of having several metrics is diagnosis. Each failure pattern points to a different component:

PatternDiagnosisFix in
Low contextual recallRetrieval missed the evidenceChunking, hybrid search, query transformation
Good recall, low contextual precisionEvidence found but ranked lowReranking
Good context, low faithfulnessModel ignores or embellishes the evidencePrompt ("answer only from context"), model choice, temperature
Faithful but low answer relevancyModel answers a neighbouring questionPrompt, query understanding, decomposition
Everything high, low correctnessLabels, or the knowledge base itself, may be wrong or outdatedData (Module 6)

LLM-as-judge

Faithfulness and relevancy cannot be computed by string matching, because they are judgements of meaning. Evaluation frameworks such as DeepEval and RAGAS use an LLM as the judge: for each test case they prompt a capable model with a rubric ("extract the claims; for each, is it supported by this context?") and turn its verdicts into a score, usually with a written reason you can read.

Practical points:

  • A test case bundles the question, the system's actual answer, the expected answer and the retrieved context. The context is needed because several metrics judge it directly.
  • Each metric has a threshold, for example 0.7, which is the pass mark for that metric's 0–1 score. It is not a required "70% word overlap" with the expected answer. A case passes the metric if its score is at or above the threshold.
  • The judge is itself a model, so it is noisy and biased: it may favour longer answers or its own phrasing. Use a strong judge model, keep the rubric fixed, and spot-check a sample of verdicts by hand, especially when a score moves unexpectedly.
  • Judging costs tokens: (number of cases) × (number of metrics) × (several LLM calls each). Run the full suite before releases, and a smaller smoke test on every change.

Building the golden dataset

Every metric needs test questions, and the correctness metrics need expected answers. Writing hundreds by hand is slow, so frameworks provide synthesizers: give them your documents and an LLM generates (question, expected answer, source context) triples, a starter golden dataset in minutes.

Synthetic goldens are a starting point, not the finish line:

  • Review them. Delete trivial or ambiguous questions and fix wrong expected answers. A judge comparing against a wrong label produces confident nonsense.
  • Add real questions. Production logs show how users actually phrase things: vague, compound, misspelt. Those are the questions that break systems.
  • Cover the hard cases deliberately: questions whose answers span several chunks, questions about codes and IDs, questions that should be answered "I don't know" because the answer isn't in the corpus, and multi-part questions.
  • Version it alongside your code, so scores remain comparable across experiments.
Evaluate like you train

Treat the golden set like a test set in machine learning. Never tune prompts on the exact questions you report scores for, or the scores will flatter you. Keep a held-out slice you look at only to confirm improvements.

EasyEvaluationFaithfulness

An answer says 'founded in 2005' while the retrieved context says 2003. Which metric catches this, and what does it tell you about where the bug is?

HardEvaluation

Contextual recall is high but answer correctness is low on a subset of questions. What could be going on?

MediumEvaluationDatasets

Why is a synthetic golden dataset not enough on its own?