Part 4

Choose the Naive Bayes variant based on the feature type.

The Bayes idea remains the same, but the likelihood model changes depending on whether features are word counts, word presence, or continuous numbers.

Three common variants

Feature question

Should repeated words matter, or only whether a word appears?

VariantFeature typeLikelihood ideaGood use case
Multinomial NBCountsHow often word/event occurs in a class.Text classification with word counts.
Bernoulli NBBinary presenceWhether feature appears or does not appear.Short text, flags, yes/no features.
Gaussian NBContinuous numbersFeature values follow class-wise Gaussian distributions.Numeric tabular features.

Multinomial vs Bernoulli

Repeated word question

Should “free free free prize” be stronger spam evidence than “free prize”?

Multinomial NB says yes because counts matter. Bernoulli NB says no after the word is already present.

MessageMultinomial viewBernoulli view
free prizefree: 1, prize: 1free: present, prize: present
free free free prizefree: 3, prize: 1free: present, prize: present
Effect of repeated words free prize free free free prize Multinomial Bernoulli

Multinomial count evidence grows with repetition; Bernoulli presence evidence does not.

Gaussian Naive Bayes

Numeric feature question

If features are petal length and petal width, how can Naive Bayes compute likelihoods?

Gaussian NB estimates a mean and variance for each feature inside each class.

\[ P(x_j\mid c)=\frac{1}{\sqrt{2\pi\sigma_{cj}^{2}}}\exp\left(-\frac{(x_j-\mu_{cj})^2}{2\sigma_{cj}^{2}}\right) \]

Then it uses the same log-score pattern:

\[ \text{score}(c)=\log P(c)+\sum_j\log P(x_j\mid c) \]

Gaussian intuition

Bell curve question

For a flower class, if most petal lengths are near 5 cm, what happens to a flower with petal length 1.5 cm?

The likelihood under that class becomes small. Gaussian NB checks how plausible each feature value is under every class distribution.

Class-wise Gaussian likelihoods new x class A class B class C

A feature value can have high likelihood under one class and low likelihood under another.

Text preprocessing choices

Preprocessing question

Should “good” and “not good” give the same evidence?

ChoiceWhy it mattersExample
LowercasingMerges duplicate forms.Free and free become same token.
Stop wordsCan reduce noise, but sometimes remove useful signals.“not” should often be kept.
N-gramsCapture short word order.“not good”, “very good”.
Binary countsReduce effect of repeated words.Useful for very short text.
Negation handlingPrevents opposite sentiment confusion.good vs NOT_good.

Evaluation

Metric question

If only 1% of messages are spam, is 99% accuracy always impressive?

No. Predicting every message as ham gives 99% accuracy but catches no spam. Use precision, recall, and F1.

\[ Precision=\frac{TP}{TP+FP} \qquad Recall=\frac{TP}{TP+FN} \]
\[ F_1=2\cdot\frac{Precision\cdot Recall}{Precision+Recall} \]
MetricIn spam detection
PrecisionOf messages predicted spam, how many were truly spam?
RecallOf all actual spam messages, how many did we catch?
F1Balanced summary of precision and recall.

Strengths and limitations

Final reflection

Why can a model with a false independence assumption still work well?

StrengthsLimitations
Very fast to train and predict.Assumes conditional independence of features.
Works well for sparse high-dimensional text.Basic bag-of-words ignores word order.
Needs relatively little data for a useful baseline.Correlated features can overcount evidence.
Interpretable through class-wise word evidence.Predicted probabilities may be poorly calibrated.
Naive Bayes is a strong baseline because it trades perfect realism for simple, stable probability estimates.
Previous: SmoothingBack to Overview