Architectures, Attention and the Transformer

2013 to 2017: deeper networks, networks that generate, networks that pay attention, and the design that took over.

A Brief History of Deep Learning

After AlexNet, progress came in a burst. Here are the ideas that shaped what followed.

Meaning as numbers: word embeddings (2013)

word2vec (Tomas Mikolov and colleagues) learned a list of numbers, a vector, for every word, just by predicting words from their neighbours. Words with similar meanings ended up near each other, and simple arithmetic on the vectors captured relationships. This showed that language can be turned into geometry that networks can work with.

Sequences, and the bottleneck (2014)

For translation, sequence-to-sequence models used one LSTM to read a sentence and compress it into a single vector, and another to write the translation. Squeezing a whole sentence into one vector loses detail on long inputs.

Attention (Bahdanau, Cho and Bengio, 2014) fixed this. While writing each output word, the model looks back over all the input words and puts more weight on the relevant ones. The idea of "look at what matters right now" turned out to be far more important than the translation task it was invented for.

Networks that create: GANs (2014)

Ian Goodfellow and colleagues proposed generative adversarial networks: one network makes fake images, another tries to tell fake from real, and each improves by competing with the other. GANs produced the first convincingly realistic generated faces.

Going deeper: ResNet (2015)

More layers should help, yet very deep networks performed worse even on their training data. Residual networks (Kaiming He and colleagues) added a shortcut that lets each block pass its input forward unchanged and learn only a correction. This made networks of 100+ layers trainable, and a 152-layer ResNet won the 2015 ImageNet contest with a top-5 error of about 3.6%, better than the roughly 5% estimated for a trained human.

A milestone in games (2016)

DeepMind's AlphaGo beat Lee Sedol, one of the world's best Go players, in 2016. Go had long been thought out of reach because the number of possible positions is enormous. It combined deep networks with search and reinforcement learning.

2017: "Attention is all you need"

Recurrent networks read a sentence one word at a time, which is slow to train and forgets over long distances. In 2017 a team at Google proposed the Transformer: drop recurrence entirely and build the network from self-attention, where every word looks directly at every other word in the sentence.

Two properties made it win:

  • It processes all words in parallel, which suits GPUs.
  • It connects distant words in one step, so long-range context is not lost.
Why this paper matters

Nearly every major language model since, and many image and audio models, is built on the Transformer. The next lesson follows what happened when people made it bigger.

MediumAttention

What problem did attention solve in sequence-to-sequence translation?

MediumTransformer

Give two reasons the Transformer replaced recurrent networks for language.