Ridge and Lasso Regression

Ridge regression adds λ‖w‖² to the squared error: the MAP estimate under a Gaussian prior, with closed form (XXᵀ + λI)⁻¹Xy, which shrinks every weight and cures collinearity. Lasso adds λ‖w‖₁ instead: a Laplace prior, no closed form, and the corners of the L1 ball push many weights to exactly zero, so it selects features.

Machine Learning Techniques

The last chapter found two problems with plain least squares. With flexible models it overfits, fitting noise with large weights. With nearly duplicate features its weights become wildly unstable. Both are cured by the same idea: penalise large weights. How you measure "large" (squared or absolute) gives two methods with strikingly different personalities.

Ridge regression

Add a penalty on the squared length of the weight vector:

min⁡w  ∑i=1n(w⊤xi−yi)2+λ∥w∥2,\min_w \; \sum_{i=1}^{n}(w^\top x_i - y_i)^2 + \lambda \lVert w \rVert^2 ,

with λ≥0\lambda \ge 0 setting how much we care about small weights. The gradient is 2X(X⊤w−y)+2λw2X(X^\top w - y) + 2\lambda w; setting it to zero gives

w^ridge=(XX⊤+λI)−1Xy.\hat w_{\text{ridge}} = \big(XX^\top + \lambda I\big)^{-1} X y .

Adding λI\lambda I raises every eigenvalue of XX⊤XX^\top by λ\lambda, so the matrix is always invertible, and the tiny eigenvalues that made least squares unstable are lifted away from zero. (In practice the intercept is usually left unpenalised.)

Why it works

  • It shrinks. In the eigenbasis of XX⊤XX^\top, ridge multiplies the least-squares component along the eigenvector with eigenvalue λj\lambda_j by λjλj+λ\frac{\lambda_j}{\lambda_j + \lambda}. Directions the data pins down well (large λj\lambda_j) are barely touched; poorly determined directions (small λj\lambda_j), which are exactly where least squares went wild, are shrunk hard.
  • Bias for variance. The ridge estimate is biased (pulled towards zero), but its variance is far smaller. Its expected squared error, bias² plus variance, can be much lower than that of the unbiased least-squares estimate. For a suitable λ\lambda it always is somewhat lower, in the sense that some λ>0\lambda > 0 beats λ=0\lambda = 0 whenever there is noise.
  • It is MAP estimation. Put a Gaussian prior on the weights, w∼N(0,τ2I)w \sim \mathcal{N}(0, \tau^2 I), and keep the Gaussian noise model. The log-posterior is
ln⁡p(w∣data)=const−12σ2∑i(yi−w⊤xi)2−12τ2∥w∥2,\ln p(w \mid \text{data}) = \text{const} - \frac{1}{2\sigma^2}\sum_i (y_i - w^\top x_i)^2 - \frac{1}{2\tau^2}\lVert w\rVert^2 ,

so the MAP estimate is ridge regression with λ=σ2/τ2\lambda = \sigma^2/\tau^2. Regularisation is a prior belief that the weights are small; the more noise relative to the prior spread, the stronger the penalty.

Lasso regression

Now penalise the sum of absolute values instead, the L1 norm:

min⁡w  ∑i=1n(w⊤xi−yi)2+λ∥w∥1,∥w∥1=∑j∣wj∣.\min_w \; \sum_{i=1}^{n}(w^\top x_i - y_i)^2 + \lambda \lVert w \rVert_1, \qquad \lVert w \rVert_1 = \sum_j |w_j| .

This is lasso (least absolute shrinkage and selection operator). In Bayesian terms it is the MAP estimate under a Laplace prior, p(wj)∝e−∣wj∣/bp(w_j) \propto e^{-|w_j|/b}, which is sharply peaked at zero.

Lasso has no closed form, because ∣wj∣|w_j| has a corner at zero where it is not differentiable. It is still a convex problem, though, and is solved efficiently by coordinate descent (update one weight at a time, holding the rest fixed) or by subgradient and proximal methods. At the corner, the role of the derivative is played by a subgradient: any slope between −1-1 and +1+1 supports the function ∣wj∣|w_j| at zero.

The striking difference: exact zeros

With one weight and an orthonormal design, the two penalties act on the least-squares value w^\hat w like this:

ridge: w^1+λ,lasso: sign⁡(w^) max⁡ ⁣(0,  ∣w^∣−λ2).\text{ridge: } \frac{\hat w}{1 + \lambda}, \qquad \text{lasso: } \operatorname{sign}(\hat w)\,\max\!\Big(0,\; |\hat w| - \frac{\lambda}{2}\Big).

Ridge scales every weight down but never to exactly zero. Lasso subtracts a fixed amount and clips at zero (soft-thresholding): any weight whose least-squares value is smaller than λ/2\lambda/2 in size becomes exactly zero. With many features, lasso therefore returns a sparse model, using only a subset of the features. It performs feature selection and fitting in one step, which is invaluable when there are many candidate features and you suspect only a few matter.

The geometric picture

Write each problem as minimising squared error subject to a budget on the weights: ∥w∥2≤t\lVert w\rVert^2 \le t for ridge, ∥w∥1≤t\lVert w\rVert_1 \le t for lasso. The contours of the squared error are ellipses centred on the least-squares solution; the solution is where the smallest ellipse first touches the budget region. The ridge region is a disc, smooth everywhere, so the touching point is almost never on an axis. The lasso region is a diamond with corners on the axes, and an ellipse expanding towards it usually hits a corner first, where some weights are exactly zero. In high dimensions the L1 ball is all corners and edges, which is why sparsity is the rule rather than the exception.

Two panels in weight space. Left: elliptical error contours around the least-squares point touching a circular ridge constraint at a point off the axes. Right: the same contours touching a diamond-shaped lasso constraint at a corner on the w2 axis, where w1 equals zero
Why lasso gives zeros. The error contours (ellipses) first touch the ridge disc at a generic point, but touch the lasso diamond at a corner, where one weight is exactly zero.

Ridge or lasso?

Ridge (L2)Lasso (L1)
Penaltyλ∥w∥2\lambda\lVert w\rVert^2λ∥w∥1\lambda\lVert w\rVert_1
PriorGaussianLaplace
Solutionclosed formiterative (coordinate descent)
Effect on weightsshrinks all, none exactly zerosets many exactly to zero
Correlated featuresspreads weight across themtends to pick one and drop the rest
Best whenmany small effects, collinearityfew features truly matter

Elastic net combines both penalties, λ1∥w∥1+λ2∥w∥2\lambda_1\lVert w\rVert_1 + \lambda_2\lVert w\rVert^2, to get sparsity while handling groups of correlated features gracefully. In every case, λ\lambda is chosen by cross-validation, and features should be standardised first, because both penalties treat all weights alike and a feature measured in large units would otherwise be penalised unfairly little.

Try it yourself
Machine Learning Lab: ridge and lasso →
Choose "Overfit", then switch on ridge and raise λ: the curve calms and test error drops. Switch to lasso and watch the weight bars: many turn grey, meaning exactly zero.
MediumRidgeLasso

With an orthonormal design, least squares gives weights (3, −0.4, 1.2). What are the ridge (λ = 1) and lasso (λ = 1) weights?

HardRidgeBayesian

Show that ridge regression is the MAP estimate under a Gaussian prior, and give λ in terms of the noise and prior variances.

MediumRidgeLassoCollinearity

Two features are almost identical copies of each other. How do least squares, ridge and lasso treat them?