Deep learning is usually told as a story about computers. It starts earlier than that, with people trying to work out what a nervous system is made of.
Is the brain one network or many cells?
In 1871 Joseph von Gerlach proposed the reticular theory: the nervous system is one continuous web, like a net with no separate pieces. Camillo Golgi, who had invented a chemical stain that made nerve tissue visible under a microscope, agreed.
Santiago Ramón y Cajal used the same stain and saw something different. He saw separate cells that touch and pass signals to each other. This became the neuron doctrine. In 1891 Wilhelm Waldeyer gave those cells the name neuron, and the evidence pushed most researchers toward Cajal.
Golgi and Cajal looked at the same kind of slide and disagreed. Data does not interpret itself. You will see this again: the same curve can look like a success or a failure depending on what you expect.
The Nobel Prize in Medicine in 1906 went to both men, which kept the quarrel alive. It took electron microscopes in the 1950s to show the tiny gap between neurons, the synapse, and settle the question in Cajal's favour.
What a neuron does, in four parts
- Dendrites receive signals from other neurons.
- The soma (cell body) combines them.
- The axon carries the result onward.
- Synapses connect to the next cell, and their strength decides how much of the signal gets through.
Keep the last point in mind. A synapse with a strength is, in later chapters, a weight. Learning in a brain changes synapse strengths, and learning in a neural network changes weights.
Local or distributed?
A second argument ran alongside the first. Is each ability, such as speech or vision, handled by one region of the brain (localized), or by many regions working together (distributed)? The answer turned out to be a mixture.
The same tension appears in today's networks. Most of a network's knowledge is spread across many weights, but researchers also build models where certain parts specialise, for example for particular languages.
Why start here
You can build neural networks without knowing any of this. But it explains the vocabulary (neuron, synapse, activation) and a pattern you will keep meeting: ideas from one field stay useful for decades and get rescued by technology from another. The neuron debate was settled by better microscopes, not by better theories.