Choose the Naive Bayes variant based on the feature type.
The Bayes idea remains the same, but the likelihood model changes depending on whether features are word counts, word presence, or continuous numbers.
Three common variants
Feature question
Should repeated words matter, or only whether a word appears?
| Variant | Feature type | Likelihood idea | Good use case |
|---|---|---|---|
| Multinomial NB | Counts | How often word/event occurs in a class. | Text classification with word counts. |
| Bernoulli NB | Binary presence | Whether feature appears or does not appear. | Short text, flags, yes/no features. |
| Gaussian NB | Continuous numbers | Feature values follow class-wise Gaussian distributions. | Numeric tabular features. |
Multinomial vs Bernoulli
Repeated word question
Should “free free free prize” be stronger spam evidence than “free prize”?
Multinomial NB says yes because counts matter. Bernoulli NB says no after the word is already present.
| Message | Multinomial view | Bernoulli view |
|---|---|---|
| free prize | free: 1, prize: 1 | free: present, prize: present |
| free free free prize | free: 3, prize: 1 | free: present, prize: present |
Multinomial count evidence grows with repetition; Bernoulli presence evidence does not.
Gaussian Naive Bayes
Numeric feature question
If features are petal length and petal width, how can Naive Bayes compute likelihoods?
Gaussian NB estimates a mean and variance for each feature inside each class.
Then it uses the same log-score pattern:
Gaussian intuition
Bell curve question
For a flower class, if most petal lengths are near 5 cm, what happens to a flower with petal length 1.5 cm?
The likelihood under that class becomes small. Gaussian NB checks how plausible each feature value is under every class distribution.
A feature value can have high likelihood under one class and low likelihood under another.
Text preprocessing choices
Preprocessing question
Should “good” and “not good” give the same evidence?
| Choice | Why it matters | Example |
|---|---|---|
| Lowercasing | Merges duplicate forms. | Free and free become same token. |
| Stop words | Can reduce noise, but sometimes remove useful signals. | “not” should often be kept. |
| N-grams | Capture short word order. | “not good”, “very good”. |
| Binary counts | Reduce effect of repeated words. | Useful for very short text. |
| Negation handling | Prevents opposite sentiment confusion. | good vs NOT_good. |
Evaluation
Metric question
If only 1% of messages are spam, is 99% accuracy always impressive?
No. Predicting every message as ham gives 99% accuracy but catches no spam. Use precision, recall, and F1.
| Metric | In spam detection |
|---|---|
| Precision | Of messages predicted spam, how many were truly spam? |
| Recall | Of all actual spam messages, how many did we catch? |
| F1 | Balanced summary of precision and recall. |
Strengths and limitations
Final reflection
Why can a model with a false independence assumption still work well?
| Strengths | Limitations |
|---|---|
| Very fast to train and predict. | Assumes conditional independence of features. |
| Works well for sparse high-dimensional text. | Basic bag-of-words ignores word order. |
| Needs relatively little data for a useful baseline. | Correlated features can overcount evidence. |
| Interpretable through class-wise word evidence. | Predicted probabilities may be poorly calibrated. |