Attention

The most-asked topic in a GenAI interview, and the one that cannot be bluffed: queries, keys and values, why scores are scaled, what masking does, how heads divide the work, and how position gets back in.

How to Crack the AI Engineer Interview

If you prepare one thing properly, prepare this. Attention is asked in almost every GenAI interview, it is asked in depth, and the follow-ups go several levels down. It is also the topic where memorised answers fall apart fastest, because the questions are mostly why.

What gets asked

The problem attention solves. Before attention, sequence models pushed everything through a fixed-size state, so distant information was squeezed or lost. Attention lets any position read directly from any other. Start here when asked to explain it — mechanism without motivation is the weaker answer.

Queries, keys and values. Each token produces three vectors. The query is what this position is looking for, the key is what each position offers, the value is what it contributes if selected. Scores come from query-key dot products, a softmax turns them into weights, and the output is the weighted sum of values. Be able to say this cleanly and then say why three projections rather than one.

Scaling by the square root of the dimension. Classic question. Dot products of high-dimensional vectors grow with dimension; large inputs push softmax into saturation, where it becomes nearly one-hot and gradients vanish. Dividing by dk\sqrt{d_k} keeps the scores in a reasonable range.

Masking. Causal masking prevents a position from attending to the future, which is what makes training parallel yet consistent with left-to-right generation. Expect the practical follow-up: what is masked, how, and why it is done with a large negative value before the softmax rather than by deletion.

Multi-head attention. Several attention operations in parallel, each in a smaller subspace, concatenated and mixed. Why: different heads specialise in different relations — syntactic, positional, topical — and one head must average those together. Know the shapes: hh heads of dimension d/hd/h, so the total cost is comparable to one full-width head.

Positional information. Self-attention is permutation-invariant: shuffle the tokens and the computation is unchanged. Position must be injected. Know the progression — integer and binary encodings, sinusoidal encodings, then rotary embeddings — and why each replaced the last.

The follow-ups that catch people

Why do queries and keys need to be different projections? Wanted: the relation is asymmetric — what a token seeks differs from what it offers — and tying them would force a symmetric score matrix.

Why is attention quadratic? Wanted: every position scores against every other, so an nn-token sequence produces n2n^2 scores, in time and memory. This sets up every efficiency question in the next chapter.

What would break without the mask? Wanted: a position could attend to tokens after it, so the model would see the answer during training and fail at generation time, where the future does not exist yet.

Why did rotary encodings replace added ones? Wanted: rotation encodes relative position directly in the attention score and extrapolates past training lengths more gracefully than a fixed table of absolute positions.

Do different heads really specialise? Wanted: yes, demonstrably — and a caution, since head interpretations are easy to over-read, which the interpretability lesson covers.

How it gets worded

This is the longest list in the course, which reflects how often attention is asked and from how many angles.

The mechanism

  • "Explain self-attention. Then tell me why we bother with three separate projections instead of using the embeddings directly."
  • "Why divide the scores by the square root of the head dimension?"
  • "What is the complexity of self-attention in sequence length, and where exactly does that come from?"
  • "Implement multi-head attention. Narrate the tensor shapes as you go."
  • "Why project and mix the heads after concatenating them?"
  • "How many heads would you use, and what are you trading when you change that number?"
  • "Do heads genuinely specialise, or is that a story we tell about them?"

Masking

  • "What is masking for? Separate a padding mask from a causal one."
  • "How does a decoder-only model avoid seeing the future while still training on a whole sequence at once?"
  • "Write the causal mask. Why add a large negative number rather than multiply by zero?"
  • "You have a padding mask and a causal mask. How do they combine?"
  • "Does the mask change between training and generation?"
  • "A model appears to be attending to padding. How do you confirm it, and how do you fix it?"
  • "BERT reads in both directions and GPT does not. Why? And how is BERT's mask token a different thing from causal masking?"

Cross-attention

  • "Self-attention and cross-attention — what changes in the computation?"
  • "In a translation model, how does the decoder decide which source words matter?"
  • "Why does a decoder-only model have no cross-attention at all?"
  • "Where does cross-attention show up outside text — in a model that takes images and language together, say?"
  • "Compare how an encoder-only, a decoder-only and an encoder-decoder model use attention."
  • "The input document is very long. What does cross-attention cost you, and what would you do about it?"

Position

  • "Shuffle the input tokens. Why does the output not change, and what do you add so that it does?"
  • "Absolute or relative position — what separates them, and when do you want each?"
  • "Sinusoidal encodings: why sine and cosine, of all things?"
  • "Learned position embeddings or fixed ones. What are you trading?"
  • "What is RoPE, and why did it take over? How does ALiBi attack the same problem differently?"
  • "A request arrives longer than anything the model was trained on. What now?"
  • "You insert a token in the middle of a sequence. What happens to every position after it?"

Reading path

From Large Language Models, in order — this is the course's own arc and it is built to be read straight through:

  1. Why Attention — the bottleneck it removes.
  2. Self-Attention — queries, keys, values from first principles.
  3. Scaled Dot-Product Attention — the scaling question, answered properly.
  4. Causal Attention — masking and why training can be parallel.
  5. Multi-Head Attention, The Shapes of Multi-Head Attention, Why Multiple Heads Help — the mechanism, the dimensions, the motivation.
  6. The Problem of Order through Rotary Positional Encodings — four short lessons building from integers to RoPE.

The Transformer Lab has an attention tab. Change a word in a sentence and watch which positions the weights move to; this is the fastest route to an intuition you can describe under pressure.

How to answer the big one

"Explain attention" is an invitation to show structure. A good two-minute answer moves in this order: the problem (fixed-size state loses distant information), the mechanism (query, key, value; scores; softmax; weighted sum), one detail that shows depth (scaling, or masking), and the cost (quadratic in sequence length) — which hands the interviewer their next question and shows you know where it leads.

Then stop. Trailing off into everything you know about transformers is a weaker answer than a tight one that invites the follow-up.