More AI engineering jobs are RAG jobs than any other kind, and interviews reflect that. Expect this topic in at least two forms: a mechanism question in the technical round, and a full system design — "build an assistant over our documents" — later on.
What gets asked
The problem it solves. A model's knowledge is fixed at training time, cannot include your private data, and when it lacks a fact it produces a fluent guess rather than nothing. Retrieval changes the task from recall to reading: find the relevant passages, put them in the prompt, answer from them. Lead with this framing — the mechanism makes more sense after it.
The pipeline. Offline: load documents, chunk them, embed the chunks, store the vectors. Online: embed the query, retrieve the nearest chunks, build a prompt around them, generate. Be able to walk both halves.
Chunking. The decision that quietly determines quality. Too small and context is lost; too large and retrieval returns mostly irrelevant text. Fixed-size with overlap, recursive splitting, structure-aware splitting on headings, and semantic splitting at topic boundaries — know the options and that the right one depends on the documents.
Embeddings and vector search. How text becomes a vector, how similarity is measured, and why exact nearest-neighbour search does not scale — hence approximate indexes, which trade a little recall for a lot of speed.
Improving retrieval. Hybrid search combining keyword and vector matching, because dense retrieval is poor at exact identifiers and rare terms. Metadata filtering. Reranking the top results with a stronger model. Rewriting the query before retrieving.
Evaluation. Retrieval and generation fail differently and need separate measurement. If the right chunk was never retrieved, no prompt fixes it.
The follow-ups that catch people
Why does RAG reduce hallucination rather than eliminate it? Wanted: the model is reading rather than recalling, which removes the main cause — but it can still misread, over-generalise, or answer from its own knowledge when retrieval returns nothing useful.
Retrieval returns nothing relevant. What happens? Wanted: an honest account. Without explicit handling, the model answers anyway from its parametric knowledge, confidently. The fix is a relevance threshold and an instruction to decline rather than guess.
Why hybrid search rather than pure vector search? Wanted: embeddings capture meaning and are weak on exact matches — product codes, error numbers, rare proper nouns. Keyword search is the opposite. Combining covers both.
How do you pick a chunk size? Wanted: by the structure of the documents and the shape of the questions, validated by measuring retrieval quality — not by a default.
Your assistant gives a wrong answer. How do you find out where it broke? Wanted: check retrieval first. Was the right passage in the retrieved set? If not, it is an indexing or retrieval problem; if yes, it is a prompt or generation problem. Candidates who start debugging the prompt have skipped the likelier cause.
How it gets worded
- "Walk me through a retrieval pipeline end to end — the offline half and the online half."
- "What does retrieval fix that a bigger context window does not?"
- "How would you chunk this set of documents, and what would tell you the chunk size is wrong?"
- "Why is exact nearest-neighbour search impractical, and what do approximate indexes give up?"
- "Why bother combining keyword search with vector search?"
- "What does a reranker add that simply retrieving more results would not?"
- "The user's question is vague. What happens before you retrieve anything?"
- "The corpus changes daily. How do you keep the index current without rebuilding it?"
- "How do you measure the retrieval half separately from the generation half?"
- "Nothing relevant comes back. What should the system do, and what does it do by default?"
- "The answer was wrong. How do you find out which stage broke?"
- "Design an assistant over our internal documents. What do you need to know before you start?"
Reading path
Advanced RAG is this chapter in full. A route sized for interview preparation:
- Why Retrieval — the problem and the pipeline.
- RAG, Fine-Tuning or Long Context — the decision interviewers love.
- Chunking and Embeddings for Retrieval — the indexing half.
- Vector Databases, Measuring Similarity, Vector Indexes — search at scale and the recall-speed trade.
- Hybrid Search, Metadata Filtering, Reranking, Query Transformation — the four upgrades that matter most.
- Keeping the Index Fresh — the operational question nobody prepares for.
- A Production RAG Blueprint — the whole thing assembled; good revision before a design round.
For depth where it comes up: GraphRAG, Multimodal RAG, and Memory for Assistants.
Handling the design round
"Build an assistant over our documentation" rewards a structure. Ask about the corpus — size, format, how often it changes — and about the bar: latency, accuracy, cost per request, and what happens when the system does not know. Then walk the pipeline, naming a choice and a reason at each stage, and say what you would measure. Finish with failure handling: stale index, empty retrieval, wrong-but-confident answers.
The two things that most reliably mark a strong answer: mentioning cost and latency unprompted, and saying how you would know it was working. Both are the marks of someone who has run one of these rather than read about it.