Why Retrieval? The Problem RAG Solves

Language models are fluent but closed-book: they cannot see private data, they do not know recent events, and they guess when unsure. RAG turns the exam into an open-book one. This chapter sets out the pipeline the rest of the course improves.

Advanced RAG Masterclass

Imagine asking a brilliant new hire about your company's notice period on their first morning. They are articulate and well read, but they have not seen your HR handbook. If they feel they must answer anyway, they will produce something that sounds like a notice-period policy, perhaps "30 days, as is common", stated with complete confidence. That is roughly the position a large language model is in whenever you ask it about something it was never trained on.

Three limits of a closed-book model

A language model learns everything it knows during training, from a large but fixed body of mostly public text. Three practical limits follow.

  1. It cannot see private data. Internal policies, Confluence pages, contracts, support tickets and design documents were never in the training data, so the model has no way of knowing what they say.
  2. Its knowledge is frozen. Training stops at a cut-off date. A policy that changed last month, a product released this week or a price updated this morning does not exist for the model.
  3. It guesses when unsure. At heart, a language model predicts a plausible next token (From Logits to Tokens shows exactly how). Plausible is not the same as true. When the model lacks a fact it still produces fluent text, and a fluent wrong answer is a hallucination. In a casual chat that is a nuisance. In a bank, a hospital or a legal team, a single confident wrong answer has a real cost.

The fix is not to make the model remember more. It is to change the exam from closed-book to open-book: find the relevant pages first, put them in front of the model, and ask it to answer from those pages.

Retrieval, augmentation, generation

The name describes the three steps exactly.

  • Retrieval: given the user's question, search your knowledge (PDFs, web pages, wikis, slides, databases) and fetch the few passages most likely to contain the answer.
  • Augmentation: build a prompt that combines the retrieved passages with the question, together with an instruction such as "answer only from the context below".
  • Generation: the language model reads that augmented prompt and writes the answer.

The model's job changes in an important way. It no longer has to know the answer. It has to read the answer from text it has been given, and summarising, extracting and rephrasing are tasks language models do very well. Retrieval also tends to bring back more text than the question needs, so the generation step serves as a filter too, keeping what is relevant and ignoring the rest.

The two halves of a RAG system

A RAG system has two phases that run at very different times.

Indexing (offline, ahead of time). You prepare the knowledge once, and again whenever it changes:

  1. Load the documents: PDFs, Word files, HTML, Markdown, slides.
  2. Split them into small passages called chunks. A whole document is too large to send to the model for every question, and too broad for search to match precisely.
  3. Embed each chunk: turn it into a vector, a list of numbers that captures its meaning, so that passages with similar meaning end up close together.
  4. Store the vectors, together with the original text and useful metadata, in a vector database.

Querying (online, per question). When a user asks something:

  1. Embed the question with the same embedding model.
  2. Search the vector database for the chunks whose vectors are closest to the question's vector, usually the top 3 to 10.
  3. Augment the prompt with those chunks.
  4. Generate the answer.
Block diagram: an offline indexing pipeline (load, chunk, embed, store) feeding a vector database, and an online query pipeline (embed question, search, augment prompt, generate answer) that reads from it
The two halves of RAG. Indexing runs ahead of time and fills the vector database; each query embeds the question, retrieves the nearest chunks and lets the model answer from them.

Meaning, not keywords

Suppose the HR policy contains a line about "10 sick-leave days annually". A user types "I am unwell, what leave can I take?" The words "sick" and "leave days" never appear in the question, yet a basic RAG system still finds the right passage. This is because the search compares meanings rather than words: the embedding of "I am unwell" sits close to the embedding of "sick leave". This is the first quiet surprise of RAG, and Module 2 explains how it works. Module 4 then covers the cases where meaning-based search is not enough, such as product codes and error IDs, and how to cover them as well.

Where a basic RAG pipeline breaks

The pipeline above works well in a demo. In production it fails in predictable places, and each failure is the subject of a later module:

Where it breaksSymptomWhere we fix it
ChunkingThe answer is split across two chunks, or buried in a huge oneModule 2
SearchThe right chunk exists but is ranked 40thModules 3 and 4
QueryThe question is vague, compound or uses different vocabularyModule 4
MeasurementNobody can tell if a change made things betterModule 5
FreshnessOld and new versions of a policy are both retrievedModule 6
Wrong sourceThe answer lives in a SQL table, not a documentModule 7
RelationshipsThe answer needs facts joined across documentsModule 8

Keep this table in mind. The rest of the course goes through it row by row.

EasyRAG fundamentals

Give three reasons a plain LLM is unsuitable as a company knowledge assistant, and explain how RAG addresses each.

MediumEmbeddingsRAG fundamentals

Why must the question be embedded with the same model that embedded the documents?

MediumRAG fundamentals

What role does the generation step play beyond 'writing the answer'?