Flip a coin three times and get three heads. Maximum likelihood says the coin lands heads with probability 1 and tails are impossible. Nobody actually believes that. You have seen many coins, and nearly all of them are close to fair; three heads in a row happens with a fair coin one time in eight. What you believed before the flips should count for something. Bayesian estimation is the formal way to let it.
The parameter as a random variable
Maximum likelihood treats the unknown parameter as a fixed number we are trying to pin down. The Bayesian view treats our uncertainty about as a probability distribution.
- The prior describes what we believe about before seeing data.
- The likelihood is the same as before: how probable the observations are for a given .
- Bayes' rule combines them into the posterior , what we believe after seeing the data:
The denominator does not depend on ; it only makes the posterior integrate to 1. So the posterior is the prior, reweighted by how well each value of explains the data.
A prior for a probability: the Beta distribution
The parameter lives between 0 and 1, so the prior must too. The standard choice is the Beta distribution with shape parameters :
where is the normalising constant. Different shapes express different beliefs.
- is flat: every bias equally plausible.
- peaks at 0.5: "probably close to fair".
- leans towards small : "probably rarely heads".
Its mean is , and the larger , the more concentrated (confident) it is.
Updating: pseudo-counts
With heads in flips, the likelihood is . Multiply by the Beta prior:
That is again a Beta distribution:
The update is just addition. The prior behaves as if we had already seen heads and tails (or, loosely, and pseudo-counts), and the data adds its real counts. When the prior and posterior belong to the same family like this, the prior is called conjugate to the likelihood. Conjugacy is what makes Bayesian updating a one-line formula here.
From a posterior to a single number
The posterior is a whole distribution, which is its strength: it says how uncertain we still are. When a single estimate is needed, two summaries are common.
- Posterior mean: .
- Posterior mode, the maximum a posteriori (MAP) estimate: .
Back to three heads in three flips, with the prior ("probably near fair"):
- MLE: .
- Posterior: , with mean and MAP .
The data has moved our belief towards heads, but three flips are not enough to overturn a reasonable prior. Now suppose we keep flipping and see 700 heads in 1,000 flips. The posterior is with mean , almost exactly the MLE of 0.7. With plenty of data, the likelihood dominates and the prior hardly matters; with little data, the prior keeps the estimate sensible.
The flat prior gives a posterior mean of , known as Laplace's rule of succession: after three heads in three flips it predicts heads with probability , not 1. The same "add one to every count" idea will return as Laplace smoothing in naive Bayes.
Maximum likelihood versus Bayesian estimation
| Maximum likelihood | Bayesian | |
|---|---|---|
| The parameter is | a fixed unknown number | a random variable with a prior |
| Uses | the likelihood only | likelihood × prior |
| Result | one point estimate | a whole posterior distribution |
| Small samples | can be extreme (e.g. ) | pulled towards the prior |
| Large samples | agrees with maximum likelihood |
The MAP estimate is the bridge between the two. It maximises : the log-likelihood plus a penalty from the prior. In regression we will see that a Gaussian prior on the weights turns least squares into ridge regression, so regularisation is, in this precise sense, a prior belief that the weights are small.