Scale and Large Language Models

2018 to today: pre-trained Transformers, scaling laws, and the path from research models to everyday assistants.

A Brief History of Deep Learning

Pre-train once, use everywhere (2018)

Before 2018 most language models were trained from scratch for each task. Two models changed that:

  • BERT (Google) learned by filling in masked-out words in huge amounts of text, then could be adapted to many tasks with little extra data.
  • GPT (OpenAI) learned by predicting the next word, which gives a model that can also write text.

This two-step recipe, pre-training on general text and then fine-tuning or prompting for a task, is how almost all large models are built today.

More scale, better results

GPT-2 (2019) and GPT-3 (2020, about 175 billion parameters) showed something striking. Making the model and training data much larger did not just improve accuracy gradually. New abilities appeared, such as following instructions given in the prompt with no extra training.

In 2020 researchers at OpenAI reported scaling laws: model quality improves predictably, following smooth curves, as you increase model size, data and compute. That made scaling a planned investment, and it is why training budgets grew so large.

Try it yourself
LLM inference simulator →

Watch a language model predict the next token, one step at a time, and see how text gets generated.

From model to assistant (2022)

A model that predicts the next word is not automatically helpful. Systems such as ChatGPT (late 2022) added a further training stage where people rated the model's answers, and the model was tuned to produce the kind people preferred (known as reinforcement learning from human feedback). Combined with a simple chat interface, this brought large language models to hundreds of millions of people within months.

Beyond language

The same ideas spread:

  • Images from text, using diffusion models, which learn to turn noise into pictures.
  • Science: DeepMind's AlphaFold2 (2020) predicted protein structures with accuracy close to laboratory methods.
  • Code, audio, video and models that handle several of these at once.

Recognition

In 2018 Geoffrey Hinton, Yoshua Bengio and Yann LeCun received the Turing Award for their work on deep learning. In 2024, John Hopfield and Hinton shared the Nobel Prize in Physics for foundational work on neural networks, and the Chemistry prize went to Demis Hassabis and John Jumper, with David Baker, for protein structure prediction and design.

The story is still open

Scale has limits, costs energy and data, and models still make confident mistakes. How far scaling carries, and what must be added beyond it, is a live research question. This lesson describes where history has reached so far, not where it will end.

EasyLLM

What are the two stages used to build most modern language models?

MediumScaling

What do scaling laws say, and why do they matter for planning?