Why Attention

The meaning of a word depends on the other words around it. Attention is the mechanism that lets each token gather, on demand, exactly the context it needs — and unlike older approaches it does so in parallel and across any distance.

Large Language Models: From Transformers to Frontier Models

Every chapter of the last module treated a token's vector as mostly self-contained. But language does not work that way. Consider:

The animal didn't cross the street because it was too tired. The animal didn't cross the street because it was too wide.

The word it points to the animal in the first sentence and the street in the second. Nothing about it itself tells you which — you have to look at tired versus wide and reach back to the right noun. A token's meaning is contextual, and a model that only ever looks at one token at a time can never resolve this. It needs a way to let it gather information from the other words. That way is attention.

What we need from the mechanism

Think about what "gathering context" really demands:

  1. Content-based, not position-based. it should find its referent by meaning ("which earlier word is a plausible thing to be tired?"), not by a fixed rule like "look three words back." The relevant word could be anywhere.
  2. Any distance, equally. The referent might be 2 tokens back or 200. The mechanism should reach far without the signal fading over distance.
  3. Parallel. For training to be efficient, every position should be able to gather its context at the same time, not one after another.
  4. A weighted blend, not a single pick. Usually a token draws a little from several places — a pronoun's meaning might combine its candidate noun, the verb, and the surrounding phrase. So we want a soft, weighted combination, not a hard choice.

Attention is precisely a mechanism with these four properties.

The older way, and why it struggled

Before attention, sequence models (recurrent networks) processed tokens one at a time, carrying a single running "memory" vector forward. This had two crippling problems for long text:

  • Distance faded. Information from far back had to survive being rewritten at every step in between, so long-range links (like it → a noun far away) were weak. This is the vanishing-signal problem you may have met in the Deep Learning course.
  • No parallelism. Because step tt needed the memory from step t−1t-1, the whole sequence had to be processed in order — slow to train on long texts.

Here is a way to feel the first problem. Read a long paragraph once, then close your eyes and translate it into another language from memory. You cannot — not faithfully — because translation needs every word, but all you kept was a blurry summary. That is exactly the bind a recurrent model is in: it must squeeze a whole passage into one fixed-size memory vector and then translate from that alone. The longer the passage, the more is lost. This is often called the context bottleneck: everything the model knows about the past has to pass through a single vector, and a single vector can only hold so much.

Attention throws out that single running memory. Instead it keeps every token's vector available at once and lets each position reach directly to any other in a single step. It is the difference between translating with your eyes closed and translating with the whole text open in front of you, free to look back at any word the moment you need it. Distance stops mattering, and because there is no left-to-right dependency in the gathering step, it all happens in parallel.

The one-line intuition

Attention lets every token ask a question about what it needs, compare that question against what every other token offers, and pull back a blend of their contents weighted by how well they match. The rest of this module is just making "question", "offer", "contents" and "match" precise.

A library analogy

Picture each token as a person in a library.

  • Each person holds up a query: "I'm a pronoun looking for a singular animate noun."
  • Every book on the shelves advertises a key: "I'm about an animal", "I'm a verb", "I'm punctuation."
  • The person compares their query to every key, giving each book a relevance score.
  • They then take a blend of the books' contents — a lot from the well-matching books, little from the rest — weighted by those scores.

That blend is the token's updated vector: itself, enriched by the context it most needed. Every person does this simultaneously, and a book can be equally useful to many readers at once.

The next chapter makes these three roles — query, key, and content (value) — into actual vectors computed from each token, and turns "compare and blend" into a precise formula.

EasyAttention

Why can't a model that processes one token at a time reliably resolve what ' it' refers to in a long sentence?

MediumAttention

Attention produces a weighted blend of other tokens' contents rather than selecting a single token. Why is 'soft' blending desirable?

MediumAttention

Name two properties attention has that a left-to-right recurrent memory lacks.