Before studying anything, it helps to know what the rounds are for. The same topic — say, fine-tuning — is a definition question in a screen, a trade-off question in a technical round, and a cost question in a design round. Knowing which conversation you are in changes the right answer.
The rounds
The screen. Usually a recruiter or a quick technical call. Breadth over depth: can you talk about what you have built, do the words mean something to you, is your stated experience real. Answers should be short and concrete.
Coding. Standard programming, sometimes with a data or ML flavour — manipulate a dataset, implement a metric, write the loop for a small training routine. Rarely about machine learning theory.
Machine learning and deep learning fundamentals. Overfitting, regularisation, optimizers, loss functions, backpropagation. Expect "explain it" followed immediately by "why does that work".
GenAI and LLM depth. The round that distinguishes an AI engineer interview from a generic ML one. Attention, tokenization, context windows, fine-tuning versus retrieval, decoding, evaluation, hallucination. This is where preparation shows most clearly.
System design. Build something: a support assistant over internal documents, a summarisation service, a code helper. Assessed on judgement rather than recall — what you choose, what you rule out, what you say about cost, latency and failure.
Behavioural. Projects you have shipped, decisions you made, things that went wrong. Specificity is everything.
What each round is really assessing
Three different things wear similar clothes.
Recall — do you know what a thing is. Necessary and insufficient. A fluent definition with no follow-through is the most common way a strong-sounding candidate falls down.
Reasoning — do you know why. Why scale attention scores by the square root of the dimension. Why a KV cache turns quadratic work into linear. Why a low-rank update is enough to adapt a model. These answers cannot be memorised convincingly, which is exactly why they are asked.
Judgement — do you know when. Given a budget, a latency target and a dataset, what would you actually build. This is the whole of the design round and most of the senior bar.
Most chapters here are aimed at the second and third. The first comes free once those are in place.
How one topic changes shape across rounds
Take retrieval-augmented generation.
- Screen: what is RAG? A sentence or two: retrieve relevant passages, put them in the prompt, have the model answer from them.
- Fundamentals: why does it reduce hallucination? Now you need the mechanism — the model is reading rather than recalling, and recall is where fabrication happens.
- GenAI depth: how do you chunk, which retrieval method, how do you evaluate it, what happens when retrieval returns nothing relevant?
- Design: here is the corpus, the latency budget and the accuracy bar — build it, and justify each choice.
One topic, four conversations. The preparation that handles all four is understanding the mechanism well enough to reason from it, which is what the linked lessons are for.
The failure modes that cost people offers
Confident guessing. The single most damaging habit. Interviewers ask follow-ups precisely to find the edge of your knowledge, and reaching it is fine — bluffing past it is not. "I have not worked with that; my guess would be X, and here is how I would check" is a strong answer.
Definitions without mechanism. Being able to say what a KV cache is, but not what it stores or why it grows, signals reading rather than understanding.
One-sided answers. Every technique in this field is a trade-off. An answer that only lists advantages invites the obvious follow-up, and not having one ready is worse than volunteering the downside yourself.
Ignoring cost. Latency, memory and money decide most real design decisions. Candidates who never mention them read as having only studied, never shipped.
Where to start
If you are unsure which area is weakest, try explaining each of these out loud, briefly: how attention works; why generation without a KV cache is slow; when you would fine-tune instead of retrieving; how you would tell whether a summariser is any good. Whichever one you cannot finish cleanly is where to begin.
The next chapter turns that into a plan.