Twenty-two chapters ago, RAG was four boxes: load, chunk, embed, retrieve. Each module since has added a layer to fix a specific failure: vague queries, codes that embeddings cannot read, the right chunk ranked fifth, stale policies, answers that live in tables, questions that span documents, users who expect to be remembered. This final chapter puts the layers together, and, as important, says which ones you actually need first.
The reference architecture
1. Ingestion (Modules 2, 6 and 10)
- Connectors to the source systems (drives, wikis, ticketing tools, databases), driven by change events where possible, so that creates, updates and deletes flow through automatically.
- Parsing that understands layout: headings, tables, figures, OCR for scans.
- Chunking that is structure-aware, with a size cap and a little overlap, sized with the embedding model's tokenizer.
- Enrichment: captions for images and summaries for tables, plus the section title prepended to each chunk.
- Metadata designed up front: source URL, title, section, department, version or effective date, access group, content hash.
- Deterministic IDs and upserts, document-level deletion of orphans, and content hashing to skip unchanged chunks.
2. Stores (Modules 2, 7, 8 and 9)
- A vector index supporting hybrid (dense plus sparse) search and filtered search, partitioned by namespace or tenant.
- Object storage for originals (figures, full tables, source files).
- A SQL database (or APIs) for structured facts, reached through text-to-SQL or dedicated tools.
- A knowledge graph, only if your questions need relationships across documents.
- A memory store of user profiles and memory collections, if the assistant should remember people.
3. The query path (Modules 3, 4 and 7)
- Route: small talk, documents, SQL, graph, web, or a clarifying question.
- Transform the query: rewrite, decompose compound questions, add step-back or HyDE where evaluation shows a gain.
- Retrieve with hybrid search, server-enforced metadata filters (tenant, access, current version) and a generous candidate pool (around 50).
- Rerank to a handful with a cross-encoder.
- Check: grade relevance, and rewrite and retry with a cap, or exit honestly.
For complex, multi-part questions, a deep agent runs this path per sub-question, calling retrieval strategies as MCP tools.
4. Generation
- Prompt assembly: system instructions ("answer only from the context; say when it isn't there"), the user's memories, the reranked evidence labelled with source numbers, and the conversation history, all within a context budget.
- A model suited to the task, vision-capable if figures are involved.
- Citations for every claim, rendered as links to the source document and page.
- Guardrails: refuse unsupported answers, mask sensitive data, and validate any structured output.
5. Observability and evaluation (Module 5), around everything
- Tracing of every request as a tree of runs, with prompts, chunks, latency and tokens, tagged by version.
- Feedback from users and reviewers attached to runs.
- A golden evaluation set, versioned, mixing synthetic, real and deliberately hard questions, and run before every release.
- Online evaluation: sampled faithfulness and relevancy judgements on live traffic.
What to build first
The full architecture is a destination, not a starting point. Every layer adds cost, latency and moving parts. A sensible order, with each step taken only when evaluation shows the previous ones are not enough:
- Baseline: good parsing, structure-aware chunking, one embedding model, top-k retrieval, a grounded prompt with citations. Build the evaluation set at the same time. Nothing after this point can be judged without it.
- Metadata and freshness: filters, versioning, deterministic IDs, deletion. These are cheap, and they prevent the most embarrassing failures.
- Hybrid search, if your users ask about codes, names or jargon.
- Reranking, if recall@50 is good but recall@5 is not.
- Query transformation and a corrective loop, if real queries are vague or compound.
- Routing to structured data, if questions are about records and numbers.
- Multimodal ingestion, if the answers sit in tables and figures.
- Memory, if personalisation matters.
- GraphRAG and deep agents, if questions genuinely need multi-hop reasoning or research-style decomposition.
Go-live checklist
- Evaluation set with real user questions, including out-of-scope questions that should get "I don't know"
- Recall@k, MRR or NDCG, faithfulness and answer relevancy above agreed thresholds
- Every answer cites its sources, and the links open
- Updates, deletions and re-indexing tested end to end, with no stale or orphan chunks
- Tenant and access filters applied by server code, and tested with adversarial prompts
- Text-to-SQL (if any) read-only, allow-listed, user-scoped, row-limited and logged
- Tracing on, with sensitive fields masked and a retention limit set
- Latency and cost per request measured at the 95th percentile, with budgets for loops and reranking
- A feedback route from users to traces to the golden set
- A rollback plan for index and model changes (blue–green re-index)
Where to go next
The field moves quickly, but the ideas in this course are stable foundations. To go deeper, read the original papers behind the techniques: dense passage retrieval, BM25, HNSW, ColBERT and ColPali, RAG-Fusion, HyDE, Self-RAG and Corrective RAG, and GraphRAG. For the model side (how the transformer that reads your context actually works, why long contexts are expensive, and how attention and the KV cache behave), see our Large Language Models course.
Above all, keep the habit this course has tried to build: every component exists to fix a measured failure. Start simple, measure honestly, read the traces, and add complexity only where the numbers say it pays off.