Causal Attention: Don't Peek Ahead

A model that predicts the next token must never see it. A triangular mask on the attention scores enforces this, and it is exactly what lets a model learn from every position of a sentence at once.

Large Language Models: From Transformers to Frontier Models

Self-attention as we built it lets every token look at every token — including the ones that come after it. For understanding a finished sentence that is fine. But a language model generates, and generation has an iron rule: when predicting the token at position tt, the model may use positions 11 through tt, and must not see position t+1t+1 or beyond. Otherwise it would be predicting a word it has already been shown — cheating that teaches it nothing.

Enforcing "look back, never forward" is the job of causal masking, and it is what makes the attention in a language model causal (also called masked self-attention).

The problem in one picture

Recall the n×nn \times n score table: entry (i,j)(i, j) is how much token ii attends to token jj. "Token ii may not look ahead" means every entry where j>ij > i — the whole region above the diagonal — must be forbidden. Token 1 may attend only to token 1; token 2 to tokens 1–2; and so on. The allowed region is a lower triangle.

An n by n attention score grid with the upper triangle (future positions) masked out, leaving a lower-triangular region of allowed attention
The causal mask. Each row i may attend only to columns j ≤ i — the lower triangle. Future positions (upper triangle) are blocked before the softmax.

How the mask works

We cannot just delete the forbidden entries — the softmax needs a full row. Instead we set each forbidden score to negative infinity before the softmax:

score(i,j)←{qi⋅kj/dk,j≤i−∞,j>i.\text{score}(i, j) \leftarrow \begin{cases} q_i \cdot k_j / \sqrt{d_k}, & j \le i \\[4pt] -\infty, & j > i. \end{cases}

Why −∞-\infty? Because softmax exponentiates, and e−∞=0e^{-\infty} = 0. A masked position therefore receives exactly zero attention weight, and the remaining (allowed) weights still sum to 1 among themselves. In practice a very large negative number stands in for −∞-\infty, but the effect is the same: the future contributes nothing, and each token's output is a blend of itself and its past only.

The payoff: train on every position at once

Masking is not just a correctness patch — it is what makes training efficient. Feed in a sentence of nn tokens. With the causal mask in place, in a single forward pass the model produces a next-token prediction at every position simultaneously:

  • position 1 predicts token 2 (having seen only token 1),
  • position 2 predicts token 3 (having seen tokens 1–2),
  • …
  • position n−1n-1 predicts token nn (having seen tokens 1 to n−1n-1).

Each of these predictions is made without any leak from the future, because the mask guarantees it. So one sentence gives n−1n-1 supervised prediction problems at once, all trained together with the ordinary cross-entropy loss. This parallel, leak-free training over every position is a large part of why transformers train so efficiently.

Training sees the whole sentence; generation grows it

During training the full target sentence is present, and the mask is what stops each position from seeing its own answer — so all positions train in parallel. During generation the future genuinely does not exist yet; the model produces one token, appends it, and runs again. The mask makes the two settings behave identically, which is exactly why a model trained this way can generate correctly.

Where this leaves us

We now have the complete attention operation used in a language model: project to queries, keys and values; score and scale; mask the future; softmax; blend the values. One such operation is a single attention head. The next module asks an obvious question — if one head learns one kind of relationship, why not run several in parallel? — and arrives at multi-head attention.

Try it yourself
Transformer Lab: the causal mask →

Toggle the mask and see the future positions disappear from every row.

EasyAttention

Why must a language model be prevented from attending to future tokens during training?

MediumAttention

Why are masked scores set to −∞ rather than 0 before the softmax?

MediumTraining

How does causal masking let a model learn from every position of a sentence in a single forward pass?