The perceptron gives a hard answer: or , with no sense of how sure it is. Often we need more. A bank approving loans, a doctor triaging patients and an advertiser bidding for clicks all need the probability of each outcome, not just a guess. Logistic regression is the standard discriminative way to get one. Despite its name, it is a classification method.
From a score to a probability
Keep the linear score and turn it into a probability with the sigmoid (logistic) function:
It squashes any real number into : , large positive scores approach 1, large negative scores approach 0, and . The model is
The decision boundary, where the probability is exactly 0.5, is : still a hyperplane. What logistic regression adds is a graded confidence that rises smoothly as a point moves away from the boundary in the direction of .
Turned around, the score is the log-odds of class 1:
Each weight therefore has a clean meaning: raising feature by one unit adds to the log-odds, multiplying the odds by . This interpretability is a large part of why logistic regression dominates medicine, economics and credit scoring.
Fitting by maximum likelihood
With labels , each label is a Bernoulli draw with probability . The likelihood of all the labels is
and the negative log-likelihood, the quantity to minimise, is the cross-entropy (or log) loss:
It charges a point little when the model gives its true label a high probability, and a great deal when the model is confidently wrong: predicting 0.01 for a point whose label is 1 costs .
The gradient
The sigmoid has a convenient derivative, . Using it, the gradient of the loss simplifies to
Each point contributes its prediction error (predicted probability minus true label) times its features: the same form as the least-squares gradient, with the sigmoid in place of the identity. Setting it to zero has no closed-form solution, because sits inside a non-linear function. But the loss is convex (its Hessian is positive semi-definite), so it has no misleading local minima: gradient descent, SGD or Newton's method all reach the global minimum.
Compare with the perceptron: it updates only on mistakes and by the full . Logistic regression updates on every point, by an amount proportional to how wrong its probability is, so points that are correct but unconvincing still nudge the boundary.
A trap: separable data
If the training data is linearly separable, the likelihood can be pushed ever closer to 1 by scaling up : doubling keeps the boundary but makes every probability more extreme, and correct, so the loss keeps falling towards zero. Gradient descent then never converges; the weights grow without bound and the model becomes absurdly confident.
The cure is regularisation, exactly as in ridge regression: minimise
which is the MAP estimate under a Gaussian prior on . It keeps the weights finite, makes the probabilities more realistic, and helps generalisation even when the data is not separable. Most software applies such a penalty by default.
Beyond two classes
For classes, give each class its own weight vector and replace the sigmoid by the softmax:
This is multinomial logistic regression (softmax regression), fitted with the same cross-entropy loss. It is also exactly the final layer of most neural network classifiers, including the next-token layer of a language model.