Four patterns worth remembering
- Ideas wait for their moment. Backpropagation, convolutional networks and attention-like ideas existed long before they changed the world. They needed data, compute or a training trick to catch up.
- Progress comes from other fields. Microscopes settled the neuron debate. Graphics cards for video games made deep learning fast. Astronomers gave us gradient descent.
- Hype and disappointment repeat. Bold predictions in 1958 were followed by a winter. Be interested in claims, but look for the benchmark and the evidence.
- Simple ideas scale. A weighted sum, a threshold and a learning rule in 1958. A weighted sum, a nonlinearity and gradient descent in today's models. Most of the gain came from doing the same simple thing with far more data and compute.
Check yourself
EasyTimeline
Order these events: AlexNet, Perceptron, Transformer, Backpropagation popularised, Minsky and Papert's book.
MediumHistory
Why did a technique that existed in 1986 only transform the field around 2012?
HardCritical thinking
A startup claims its new model will soon replace all human experts. Using this history, how would you evaluate the claim?
Where next
You now have the map. To build the machinery itself, continue with the Deep Learning course and its simulator, and try the Deep Learning simulator directly.
For deeper reading on this history, Jürgen Schmidhuber's survey "Deep Learning in Neural Networks: An Overview" (2015) and the 2015 Nature review "Deep learning" by LeCun, Bengio and Hinton are good places to start.