Naive Bayes on words used coin-flip features: a word is present or not. Many problems have real-valued features instead: a flower's petal length, a patient's blood pressure, a sensor's temperature. The generative recipe still applies; only the model of each feature changes. The obvious choice for a continuous measurement is a Gaussian.
The model
For a binary problem with features :
- The class prior is .
- Within class , feature is Gaussian with its own mean and variance: .
- The naive assumption: within a class, the features are independent, so
Geometrically, each class is modelled as an axis-aligned Gaussian blob: an ellipse whose axes are parallel to the feature axes, because independence means no correlation between features within a class.
Training
Each parameter has the maximum likelihood estimate from Module 5, computed on that class's points only:
where is the number of training points in class . Training is again a single pass: per class and per feature, a mean and a variance. (A small constant is usually added to each variance so that a feature that happens to be constant within a class does not get variance zero.)
Classifying, and the shape of the boundary
Compare the log posteriors. For class ,
Predict class 1 when . The boundary is where the two scores are equal. Each score contains terms in with coefficient , so in general:
The boundary is quadratic. If the two classes have different variances for a feature, the terms do not cancel, and the boundary is a curve: an ellipse, a parabola or a hyperbola in two dimensions. A tight class surrounded by a widely spread one gets a closed boundary around it.
With shared variances, the boundary is linear. If for every feature, the terms are identical in both scores and cancel, leaving
a linear classifier, with each feature weighted by how far apart the class means are relative to that feature's spread. This is the same family as logistic regression, and the same family as naive Bayes on binary features.
Generative classifiers, compared
Naive Bayes is the simplest member of a family of Gaussian generative classifiers. They differ only in how much structure the class covariance matrices may have.
| Model | Class covariance | Boundary |
|---|---|---|
| Gaussian naive Bayes, shared variances | diagonal, the same for both classes | linear |
| Gaussian naive Bayes | diagonal, per class | quadratic |
| Linear discriminant analysis (LDA) | full, shared | linear |
| Quadratic discriminant analysis (QDA) | full, per class | quadratic |
More freedom means more parameters to estimate (a full covariance matrix has entries), so the richer models need more data. Naive Bayes sits at the cheap end: it ignores correlations between features, and in return needs only numbers per class.
When the assumptions fail
Gaussian naive Bayes is fast, needs little data and gives sensible probabilities when its assumptions roughly hold. Two failures are common.
- Non-Gaussian features. A feature with two peaks within a class, or heavy skew, is poorly described by one Gaussian. Transforming the feature (a logarithm for skewed positive quantities) or using a mixture per class helps.
- Strongly correlated features. Independence counts correlated evidence twice. The predicted class may still be right, but the probabilities become overconfident, pushed towards 0 and 1.
In both cases a discriminative method such as logistic regression, which makes no assumption about how the features are distributed, is a natural alternative, and it is the next chapter's subject.