A language model knows a great deal, but it does not know your documents. It has never read your company's leave policy, last quarter's incident reports or the contract you signed on Tuesday. Retrieval-Augmented Generation (RAG) closes that gap. Before the model answers, a retrieval system finds the passages that matter and places them in front of it, so the answer is grounded in your sources rather than in the model's memory.
A basic RAG demo takes an afternoon to build. Making it dependable is much harder. Real users ask vague questions, use words your documents never use, ask three things at once, or ask about an error code that no embedding model has ever seen. Documents change every week. Some answers need a database lookup, not a paragraph of text. Some need facts joined across five different files. And an engineer has to be able to prove the system is getting better rather than just feeling that it is.
This course is about that gap between a demo and a production system. We build the standard pipeline carefully and then improve it one layer at a time. Each chapter introduces one idea, explains why it is needed, and shows where it sits in the pipeline.
Who this course is for
The course is written for postgraduate students and working engineers. It is concept-first: there is no code, only the ideas, the trade-offs and the block diagrams you would draw on a whiteboard while designing a system. The tools named along the way (LangChain, LangGraph, Pinecone, Weaviate, Milvus, Neo4j, DeepEval and others) appear only as examples. Every idea carries over to whichever stack you use.
What you should already know
- What a language model does: it predicts the next token, and it can be prompted. Our Large Language Models course covers this from scratch.
- What an embedding is, at least roughly: a vector that represents meaning. Embeddings: Tokens into Vectors is a good refresher.
- Basic linear algebra: vectors, dot products and lengths.
Alongside the chapters, the RAG Lab lets you try the main ideas hands-on in your browser: chunking, embedding search, vector indexes, hybrid search, reranking, retrieval metrics and index freshness. The chapters link to the matching experiment.
How the course is organised
- Foundations: why retrieval is needed, and when RAG is the right tool rather than fine-tuning or a long context window.
- Indexing: from documents to vectors: chunking, embeddings and vector databases.
- Search at scale: how similarity is measured, and the index structures (IVF, HNSW) that make search over millions of vectors fast.
- Sharper retrieval: hybrid search, metadata filtering, reranking and query transformation.
- Evaluating RAG: retrieval metrics, end-to-end evaluation of answers, and tracing a pipeline to find out where it fails.
- Keeping knowledge fresh: updating a live index without duplicates or stale answers.
- Agentic RAG: self-correcting retrieval loops, routing between data sources, text-to-SQL, and retrieval as tools for an agent.
- GraphRAG: when relationships matter more than similarity.
- Memory: letting an assistant remember users across conversations.
- Beyond text, into production: multimodal documents, and a blueprint that puts every piece together.
Every chapter ends with interview-style questions that check your understanding. Most of them are the kind of question asked in applied-AI interviews.
This course follows the open lecture series "RAG – Retrieval Augmented Generation" by Praveen Reddy Learnings, which goes from a first RAG chatbot to production agents. We have re-taught the material in our own words as concepts only, corrected a few slips in the formulas and worked examples, and added topics such as query transformation that complete the picture. The credit for the original sequence belongs to the series author.