A Study Plan

A four-week route through the material, what to do if you have less time, and how to study so that the knowledge survives a follow-up question.

How to Crack the AI Engineer Interview

The material below is a lot, and most of it is already written. What follows is an order to read it in and a way to read it that holds up under questioning.

Four weeks, part time

Week 1 — transformer internals. The hardest to fake and the most asked. Work through Large Language Models modules 1 to 5: what a language model is, tokenization and embeddings, the block, attention from first principles, multi-head attention, positional encodings. Do not skim attention. Everything later depends on it.

Week 2 — efficiency and adaptation. Modules 6 to 10 of the same course: the KV cache and its variants, mixture-of-experts, quantization, pruning and distillation, then alignment, fine-tuning and synthetic data. This is the week that teaches you to talk about cost, which is what separates candidates who have deployed something.

Week 3 — systems. Advanced RAG. Prompting and in-context learning, then retrieval end to end, then agents. Most AI engineer roles are mainly this, whatever the interview spends its time on.

Week 4 — evaluation, failure modes, and revision. Modules 11 and 12 of the LLM course, the evaluation module of the RAG course, and then go back over anything from weeks 1 to 3 you cannot explain without notes.

Foundations — Machine Learning Techniques and Deep Learning — sit underneath all of this. If they are solid, dip in as needed. If they are not, add a week at the front; the transformer material is much harder without gradients and backpropagation in place.

If you have less time

One week. Attention (why attention through multi-head attention), the KV cache, RAG versus fine-tuning, fine-tuning and LoRA, and hallucinations. That is the set most likely to come up.

Two days. Attention and the KV cache, and be honest about the rest.

How to study this

Explain it out loud, without notes. Reading creates a feeling of understanding that collapses the moment you have to produce the explanation yourself. If you cannot say it aloud in two minutes, you do not have it yet. This is the single highest-return habit for an interview, because it rehearses the actual task.

Follow the why chain down three levels. Take any fact and ask why until you hit something you cannot answer. Attention scores are scaled by the square root of the dimension — why? Because the dot products grow with dimension — why does that matter? Because softmax saturates on large inputs and the gradients vanish. Three levels is about as deep as an interview goes.

Learn each topic's trade-off. For every technique, know what it costs. Quantization trades precision for memory. LoRA trades expressiveness for cost. RAG trades latency for freshness. Agents trade reliability for capability. An answer that names the trade-off before being asked for it reads as experience.

Work the numbers once. Estimate a KV cache for a real model shape. Count LoRA's parameters against a full fine-tune. Reason about what a 4-bit model saves. Doing this once makes cost arguments concrete rather than vague, and the arithmetic is small.

Rehearse against the wordings, not the headings. Each chapter from here on has a section called How it gets worded, listing the sentences its topic turns up in. Work through one chapter's list at the end of the day you studied it, answering aloud and timing yourself to two minutes. The list is deliberately longer than the chapter's own follow-ups, because the point is not new material — it is to stop a familiar idea from sounding unfamiliar because it was phrased differently.

Use the simulators. Several lessons link to the Transformer Lab and the Deep Learning simulator. Watching attention weights change as you edit a sentence does more for intuition than another pass through the text.

What to prepare that is not on this syllabus

Two or three projects you can discuss properly. What you built, what you chose, what went wrong, what you measured. Interviews probe depth here, and a small project understood thoroughly beats a large one described vaguely.

Your own numbers. If you have deployed anything: latency, cost per request, throughput, what the bottleneck was. Almost nobody brings these, and they are disproportionately convincing.

Questions for them. How they evaluate models, what their inference stack looks like, how they handle hallucination in production. Good questions double as evidence that you know what the hard parts are.

The rest of this course walks each area in turn. Read quickly first, then go back to whatever you could not explain aloud.