Tokens, Embeddings and the Block

What a language model is, how text becomes numbers, and the structure of the block that everything else repeats — the groundwork the attention questions assume.

How to Crack the AI Engineer Interview

The GenAI depth round usually starts somewhere in this chapter. The questions look basic and are not: candidates who have only read about attention often cannot say what a token actually is, why a vocabulary is built the way it is, or what the feed-forward layer contributes.

What gets asked

What a language model is. A next-token predictor over a vocabulary, trained by maximising the likelihood of real text. Simple to say, and worth saying precisely, because almost every later answer refers back to it — hallucination, decoding, perplexity, alignment.

Tokenization. Why text is split into subwords rather than characters or words. The trade-off: a word-level vocabulary cannot handle unseen words, a character-level one makes sequences very long, subwords sit between. Expect questions about rare words splitting into pieces, non-English text costing more tokens, and why token counts determine cost.

Embeddings. How a token id becomes a vector, what the embedding matrix is, and that these vectors are learned rather than assigned. A common follow-up is the difference between a static embedding table and the contextual representation that comes out of the stack.

The block. Two sub-layers — attention to move information between positions, a feed-forward network to transform each position — wrapped in residual connections and normalisation. Know what each part is for. The most-missed question is what the feed-forward layer does, and the useful answer is that it holds most of the model's parameters and does most of the per-position transformation, which is precisely why mixture-of-experts targets it.

Shapes. Being able to say what dimensions a sequence has as it moves through — tokens to vectors to an n×dn \times d matrix — is a quick signal that you have actually worked with this.

The follow-ups that catch people

Why subword tokenization rather than words? Wanted: unbounded vocabulary and unknown-word failures versus sequence length, with subwords as the compromise that keeps the vocabulary fixed and still represents anything.

What is in the residual stream? Wanted: a running representation per position that each block adds a correction to, which is both why gradients flow and why blocks can be shallow adjustments rather than full rebuilds.

Why normalise before each sub-layer? Wanted: keeping activations in a stable range regardless of depth, with pre-norm being the arrangement that trains most stably.

Where do a model's parameters actually live? Wanted: mostly the feed-forward layers, not attention. People guess attention and it is the wrong answer.

How it gets worded

Three clusters: how text becomes tokens, how tokens become vectors, and what the block does with them.

  • "Take me through byte-pair encoding one merge at a time."
  • "What does splitting into subwords fix that a word-level vocabulary cannot?"
  • "A user types a word that appeared nowhere in training. What does the model do with it?"
  • "Why do GPT-style models work over bytes rather than characters?"
  • "What is a merge rule, and how is it applied when you tokenize a string you have never seen?"
  • "Why is vocabulary size worth arguing about?"
  • "Here is a corpus. Build me a tokenizer from nothing — what are the steps?"
  • "Explain an embedding to someone with no technical background. Now explain it to me."
  • "What is the difference between a fixed embedding table and the vectors coming out of the stack?"
  • "Why do dense vectors beat keyword weighting for semantic search — and when is a sparse representation still the better tool?"
  • "What is cosine similarity actually measuring in that space?"
  • "How would you produce embeddings for images or audio? What stays the same?"
  • "How do you tell whether a set of embeddings is any good?"
  • "Where do embeddings sit in a retrieval pipeline?"
  • "Write out layer normalisation. What are its learnable parameters for, and what would removing them cost?"
  • "Why do transformers normalise across features rather than across the batch? What breaks if you try batch statistics while generating one token at a time?"
  • "Pre-norm or post-norm — what moved, and why do modern decoder stacks use the first?"
  • "Delete every normalisation layer. What happens to training?"
  • "What shape is a batch of text at each stage of the block?"

Reading path

From Large Language Models, in order:

  1. What a Language Model Is — the objective, and what follows from it.
  2. How a Modern LLM Is Built — the map of the whole course in one lesson.
  3. Tokenization — subwords, vocabularies, and why token counts are costs.
  4. Embeddings — ids to vectors, and what the table learns.
  5. Inside a Transformer Block — attention, feed-forward, residual, norm.
  6. From Logits to Tokens — how the final vector becomes an actual word.

The Transformer Lab has tabs for tokenization and embeddings. Running a few sentences through the tokenizer and watching where the splits land is five minutes well spent — especially with a non-English sentence, which makes the cost asymmetry obvious.

Worth knowing beyond the basics

Context length is a cost, not a setting. Longer contexts mean more positions to attend over and a larger cache at inference. When an interviewer asks why long-context requests are priced higher, the answer is in the KV cache lesson, and it is quantitative.

Tokenization breaks comparisons. Two models with different tokenizers cannot have their perplexities compared, because the metric is averaged over tokens and the tokens are not the same. This is a favourite trap.

The block is repeated, not varied. The stack is one design, repeated with separate weights. Candidates sometimes imagine a pipeline of differently-shaped stages; it is the same block throughout.