Bayesian Estimation

Bayesian estimation treats the unknown parameter as a random variable with a prior distribution and updates it to a posterior with Bayes' rule. For a coin, a Beta prior gives a Beta posterior that simply adds the observed heads and tails to pseudo-counts. The posterior mean or mode smooths the maximum likelihood estimate and avoids its overconfidence on small samples.

Machine Learning Techniques

Flip a coin three times and get three heads. Maximum likelihood says the coin lands heads with probability 1 and tails are impossible. Nobody actually believes that. You have seen many coins, and nearly all of them are close to fair; three heads in a row happens with a fair coin one time in eight. What you believed before the flips should count for something. Bayesian estimation is the formal way to let it.

The parameter as a random variable

Maximum likelihood treats the unknown parameter pp as a fixed number we are trying to pin down. The Bayesian view treats our uncertainty about pp as a probability distribution.

  • The prior P(p)P(p) describes what we believe about pp before seeing data.
  • The likelihood P(data∣p)P(\text{data} \mid p) is the same as before: how probable the observations are for a given pp.
  • Bayes' rule combines them into the posterior P(p∣data)P(p \mid \text{data}), what we believe after seeing the data:
P(p∣data)=P(data∣p) P(p)P(data)  ∝  P(data∣p)⏟likelihood  P(p)⏟prior.P(p \mid \text{data}) = \frac{P(\text{data} \mid p)\, P(p)}{P(\text{data})} \;\propto\; \underbrace{P(\text{data} \mid p)}_{\text{likelihood}} \; \underbrace{P(p)}_{\text{prior}} .

The denominator does not depend on pp; it only makes the posterior integrate to 1. So the posterior is the prior, reweighted by how well each value of pp explains the data.

A prior for a probability: the Beta distribution

The parameter pp lives between 0 and 1, so the prior must too. The standard choice is the Beta distribution with shape parameters α,β>0\alpha, \beta > 0:

P(p)=pα−1(1−p)β−1B(α,β),0≤p≤1,P(p) = \frac{p^{\alpha - 1}(1 - p)^{\beta - 1}}{B(\alpha, \beta)}, \qquad 0 \le p \le 1,

where B(α,β)B(\alpha, \beta) is the normalising constant. Different shapes express different beliefs.

  • Beta(1,1)\text{Beta}(1, 1) is flat: every bias equally plausible.
  • Beta(10,10)\text{Beta}(10, 10) peaks at 0.5: "probably close to fair".
  • Beta(2,8)\text{Beta}(2, 8) leans towards small pp: "probably rarely heads".

Its mean is αα+β\tfrac{\alpha}{\alpha + \beta}, and the larger α+β\alpha + \beta, the more concentrated (confident) it is.

Updating: pseudo-counts

With kk heads in nn flips, the likelihood is pk(1−p)n−kp^k(1 - p)^{n - k}. Multiply by the Beta prior:

P(p∣data)  ∝  pk(1−p)n−k⋅pα−1(1−p)β−1=p k+α−1(1−p) n−k+β−1.P(p \mid \text{data}) \;\propto\; p^{k}(1 - p)^{n - k} \cdot p^{\alpha - 1}(1 - p)^{\beta - 1} = p^{\,k + \alpha - 1}(1 - p)^{\,n - k + \beta - 1} .

That is again a Beta distribution:

p∣data  ∼  Beta(α+k,  β+n−k).p \mid \text{data} \;\sim\; \text{Beta}\big(\alpha + k,\; \beta + n - k\big).

The update is just addition. The prior behaves as if we had already seen α−1\alpha - 1 heads and β−1\beta - 1 tails (or, loosely, α\alpha and β\beta pseudo-counts), and the data adds its real counts. When the prior and posterior belong to the same family like this, the prior is called conjugate to the likelihood. Conjugacy is what makes Bayesian updating a one-line formula here.

Three curves on p from 0 to 1: a Beta(5,5) prior centred at 0.5, a likelihood for 3 heads in 3 flips rising towards 1, and the Beta(8,5) posterior peaking near 0.64, between them
Bayesian updating for a coin. A prior centred on fair, Beta(5, 5), meets the likelihood of three heads in three flips; the posterior Beta(8, 5) moves towards heads but stays far from the maximum likelihood answer p = 1.

From a posterior to a single number

The posterior is a whole distribution, which is its strength: it says how uncertain we still are. When a single estimate is needed, two summaries are common.

  • Posterior mean: p^=α+kα+β+n\displaystyle \hat p = \frac{\alpha + k}{\alpha + \beta + n}.
  • Posterior mode, the maximum a posteriori (MAP) estimate: p^MAP=α+k−1α+β+n−2\displaystyle \hat p_{\text{MAP}} = \frac{\alpha + k - 1}{\alpha + \beta + n - 2}.

Back to three heads in three flips, with the prior Beta(5,5)\text{Beta}(5, 5) ("probably near fair"):

  • MLE: 3/3=13/3 = 1.
  • Posterior: Beta(8,5)\text{Beta}(8, 5), with mean 8/13≈0.628/13 \approx 0.62 and MAP 7/11≈0.647/11 \approx 0.64.

The data has moved our belief towards heads, but three flips are not enough to overturn a reasonable prior. Now suppose we keep flipping and see 700 heads in 1,000 flips. The posterior is Beta(705,305)\text{Beta}(705, 305) with mean 705/1010≈0.698705/1010 \approx 0.698, almost exactly the MLE of 0.7. With plenty of data, the likelihood dominates and the prior hardly matters; with little data, the prior keeps the estimate sensible.

The flat prior Beta(1,1)\text{Beta}(1, 1) gives a posterior mean of k+1n+2\tfrac{k + 1}{n + 2}, known as Laplace's rule of succession: after three heads in three flips it predicts heads with probability 4/54/5, not 1. The same "add one to every count" idea will return as Laplace smoothing in naive Bayes.

Maximum likelihood versus Bayesian estimation

Maximum likelihoodBayesian
The parameter isa fixed unknown numbera random variable with a prior
Usesthe likelihood onlylikelihood × prior
Resultone point estimatea whole posterior distribution
Small samplescan be extreme (e.g. p^=1\hat p = 1)pulled towards the prior
Large samplesagrees with maximum likelihood

The MAP estimate is the bridge between the two. It maximises ln⁡P(data∣p)+ln⁡P(p)\ln P(\text{data} \mid p) + \ln P(p): the log-likelihood plus a penalty from the prior. In regression we will see that a Gaussian prior on the weights turns least squares into ridge regression, so regularisation is, in this precise sense, a prior belief that the weights are small.

EasyBayesian estimation

With a Beta(2, 2) prior, you observe 1 head in 4 flips. What are the posterior, its mean and the MAP estimate? Compare with the MLE.

MediumBayesian estimationConjugacy

What does it mean for a prior to be conjugate, and why is it convenient?

MediumBayesian estimationMLE

When would you prefer the Bayesian estimate to the maximum likelihood estimate in practice?