Smoothing and log-space make Naive Bayes usable.
The raw formula is intuitive, but two practical issues appear immediately: zero probabilities and tiny multiplied numbers.
Zero-probability problem
Failure question
If one word never appeared in ham during training, should the ham score become exactly zero?
Without smoothing, one unseen word can destroy the full class score because probabilities are multiplied.
That is too harsh. Unseen in a small training dataset should not mean impossible.
Multiplication makes zero probabilities dangerous.
Laplace smoothing
Repair question
How can we give every known vocabulary word a tiny non-zero chance?
Laplace smoothing adds a small imaginary count to every word.
For \(\alpha=1\), every word gets one extra count. The denominator also increases by \(|V|\), so the probabilities still sum to 1.
| Quantity | Without smoothing | With \(\alpha=1\) |
|---|---|---|
| Word count | \(count(w,c)\) | \(count(w,c)+1\) |
| Total count | \(N_c\) | \(N_c+|V|\) |
| Unseen word | 0 probability | Small non-zero probability |
Effect of smoothing
Before-after question
Does smoothing remove the evidence from strong words like “free”?
No. Smoothing softens extreme probabilities. Strong words remain strong, but unseen words no longer create impossible classes.
| Word | \(P(w\mid ham)\) | \(P(w\mid spam)\) | Direction |
|---|---|---|---|
| free | small | large | spam |
| meeting | large | small | ham |
| today | medium | medium | weak |
Smoothing does not flatten everything; it prevents absolute zeros.
Why log probabilities?
Tiny number question
What happens when a long document multiplies 200 small probabilities?
The product becomes extremely tiny. Computers may round it to zero. Logs solve this by turning multiplication into addition.
Since log is an increasing function, the winning class does not change.
Working example: predict a new sentence
Prediction walkthrough
A new sentence comes in: free prize today. How exactly does Naive Bayes decide between ham and spam?
Assume the model has already learned these class-wise word counts from a tiny training dataset.
| Word | Count in ham | Count in spam |
|---|---|---|
| free | 0 | 4 |
| prize | 0 | 3 |
| today | 2 | 1 |
| meeting | 4 | 0 |
| project | 3 | 0 |
| report | 3 | 0 |
Let the vocabulary size be \(|V|=6\), and use Laplace smoothing with \(\alpha=1\).
The total word counts are:
So the smoothed denominators become:
| Word in new sentence | \(P(w\mid ham)\) | \(P(w\mid spam)\) |
|---|---|---|
| free | \(\frac{0+1}{18}=\frac{1}{18}\) | \(\frac{4+1}{14}=\frac{5}{14}\) |
| prize | \(\frac{0+1}{18}=\frac{1}{18}\) | \(\frac{3+1}{14}=\frac{4}{14}\) |
| today | \(\frac{2+1}{18}=\frac{3}{18}\) | \(\frac{1+1}{14}=\frac{2}{14}\) |
Assume both classes are equally common in the training data:
Now compute the two Naive Bayes scores.
The spam score is larger, so the prediction is spam.
Smoothing gives ham a non-zero score, but the sentence still has much stronger spam evidence.
In log-space, the same comparison becomes addition instead of multiplication.
Less negative is larger, so the log-space prediction is also spam.
Class imbalance: will Naive Bayes always predict ham?
Imbalance question
If most emails are ham, does the model simply predict every new email as ham?
No. A larger ham prior gives ham a head start, but strong spam-like words can overcome that head start.
For spam vs ham, compare the log-score difference:
If \(\Delta > 0\), spam wins. If \(\Delta < 0\), ham wins.
Suppose the training data is imbalanced:
The spam class starts with a disadvantage:
Now consider this new message:
Assume these words provide the following spam evidence:
| Word | Spam evidence \(\log \frac{P(w\mid spam)}{P(w\mid ham)}\) |
|---|---|
| free | 1.8 |
| lottery | 2.2 |
| prize | 1.7 |
| claim | 1.5 |
Total word evidence:
Final spam-vs-ham evidence:
Since \(4.26 > 0\), the model predicts spam even though spam is the minority class.
The prior is only the starting belief. The message words can overturn it.
Word evidence in log-space
Interpretability question
Which word in a message contributes most toward spam?
Compare a word’s log likelihood under spam and ham.
| Value | Meaning |
|---|---|
| Positive | The word pushes toward spam. |
| Negative | The word pushes toward ham. |
| Near zero | The word is not strongly class-specific. |
Unknown words at prediction time
New word question
If the model never saw the word “crypto” during training, what should happen during prediction?
A word outside the training vocabulary has no learned count. Common practical choices are:
| Approach | Meaning | When it is useful |
|---|---|---|
| Ignore unknown word | The word contributes no class evidence. | Common with fixed vocabularies such as CountVectorizer. |
| Use an UNK token | Rare or unseen words map to one shared unknown token. | Useful when unknown-word behavior itself carries information. |
| Use character n-grams | Represent pieces like prefixes, suffixes, or character windows. | Useful for spelling variants, names, and noisy text. |
Why Naive Bayes is linear in log-space
Boundary question
If each word adds or subtracts evidence, what shape does the final decision take?
For two classes, compare the spam log score and ham log score. The difference can be written as a weighted sum of word counts.
Here \(x_j\) is the count of word \(j\). This looks like a linear model on bag-of-words features: each word has a fixed evidence weight, and the document adds them up.
Full training algorithm
Model storage question
After training, what does a Naive Bayes model actually store?
- Build the vocabulary from training documents.
- Count documents per class to estimate \(P(c)\).
- Count words inside each class.
- Apply Laplace smoothing to estimate \(P(w\mid c)\).
- Store \(\log P(c)\) and \(\log P(w\mid c)\).
classes,
vocabulary,
class_log_priors,
word_log_likelihoods
}
Full prediction algorithm
Prediction question
How does the model score a new sentence step by step?
- Tokenize the new document.
- For each class, start with \(\log P(c)\).
- Add \(\log P(w\mid c)\) for every word in the document.
- Return the class with the largest log score.