Part 3

Smoothing and log-space make Naive Bayes usable.

The raw formula is intuitive, but two practical issues appear immediately: zero probabilities and tiny multiplied numbers.

Zero-probability problem

Failure question

If one word never appeared in ham during training, should the ham score become exactly zero?

Without smoothing, one unseen word can destroy the full class score because probabilities are multiplied.

\[ 0.5\times0.12\times0\times0.08=0 \]

That is too harsh. Unseen in a small training dataset should not mean impossible.

One zero collapses the product 0.5 × 0.12 × 0 = 0 A missing word should reduce confidence, not erase a class.

Multiplication makes zero probabilities dangerous.

Laplace smoothing

Repair question

How can we give every known vocabulary word a tiny non-zero chance?

Laplace smoothing adds a small imaginary count to every word.

\[ P(w\mid c)=\frac{\text{count}(w,c)+\alpha}{\sum_{w'\in V}\text{count}(w',c)+\alpha|V|} \]

For \(\alpha=1\), every word gets one extra count. The denominator also increases by \(|V|\), so the probabilities still sum to 1.

QuantityWithout smoothingWith \(\alpha=1\)
Word count\(count(w,c)\)\(count(w,c)+1\)
Total count\(N_c\)\(N_c+|V|\)
Unseen word0 probabilitySmall non-zero probability

Effect of smoothing

Before-after question

Does smoothing remove the evidence from strong words like “free”?

No. Smoothing softens extreme probabilities. Strong words remain strong, but unseen words no longer create impossible classes.

Word\(P(w\mid ham)\)\(P(w\mid spam)\)Direction
freesmalllargespam
meetinglargesmallham
todaymediummediumweak
Smoothed word likelihoods freemeetingtoday ham spam

Smoothing does not flatten everything; it prevents absolute zeros.

Why log probabilities?

Tiny number question

What happens when a long document multiplies 200 small probabilities?

The product becomes extremely tiny. Computers may round it to zero. Logs solve this by turning multiplication into addition.

\[ \log\left(P(c)\prod_{i=1}^{n}P(w_i\mid c)\right)=\log P(c)+\sum_{i=1}^{n}\log P(w_i\mid c) \]

Since log is an increasing function, the winning class does not change.

\[ \arg\max_c P(c)\prod_iP(w_i\mid c)=\arg\max_c\left[\log P(c)+\sum_i\log P(w_i\mid c)\right] \]

Working example: predict a new sentence

Prediction walkthrough

A new sentence comes in: free prize today. How exactly does Naive Bayes decide between ham and spam?

Assume the model has already learned these class-wise word counts from a tiny training dataset.

WordCount in hamCount in spam
free04
prize03
today21
meeting40
project30
report30

Let the vocabulary size be \(|V|=6\), and use Laplace smoothing with \(\alpha=1\).

\[ P(w\mid c)=\frac{\text{count}(w,c)+1}{\text{total word count in class }c+|V|} \]

The total word counts are:

\[ N_{ham}=12 \qquad N_{spam}=8 \]

So the smoothed denominators become:

\[ N_{ham}+|V|=12+6=18 \qquad N_{spam}+|V|=8+6=14 \]
Word in new sentence\(P(w\mid ham)\)\(P(w\mid spam)\)
free\(\frac{0+1}{18}=\frac{1}{18}\)\(\frac{4+1}{14}=\frac{5}{14}\)
prize\(\frac{0+1}{18}=\frac{1}{18}\)\(\frac{3+1}{14}=\frac{4}{14}\)
today\(\frac{2+1}{18}=\frac{3}{18}\)\(\frac{1+1}{14}=\frac{2}{14}\)

Assume both classes are equally common in the training data:

\[ P(ham)=P(spam)=0.5 \]

Now compute the two Naive Bayes scores.

\[ score(ham)=0.5\times\frac{1}{18}\times\frac{1}{18}\times\frac{3}{18}=0.000257 \]
\[ score(spam)=0.5\times\frac{5}{14}\times\frac{4}{14}\times\frac{2}{14}=0.007289 \]

The spam score is larger, so the prediction is spam.

\[ \hat{c}=spam \]
New sentence: free prize today free prize today ham score 0.5 x 1/18 x 1/18 x 3/18 = 0.000257 spam score 0.5 x 5/14 x 4/14 x 2/14 = 0.007289 larger score wins: spam

Smoothing gives ham a non-zero score, but the sentence still has much stronger spam evidence.

In log-space, the same comparison becomes addition instead of multiplication.

\[ \log score(ham)=\log(0.5)+\log\frac{1}{18}+\log\frac{1}{18}+\log\frac{3}{18}=-8.266 \]
\[ \log score(spam)=\log(0.5)+\log\frac{5}{14}+\log\frac{4}{14}+\log\frac{2}{14}=-4.922 \]

Less negative is larger, so the log-space prediction is also spam.

This example shows the full prediction mechanics: tokenize the sentence, look up smoothed likelihoods, multiply or add log probabilities, then choose the class with the larger score.

Class imbalance: will Naive Bayes always predict ham?

Imbalance question

If most emails are ham, does the model simply predict every new email as ham?

No. A larger ham prior gives ham a head start, but strong spam-like words can overcome that head start.

For spam vs ham, compare the log-score difference:

\[ \Delta = score(spam)-score(ham) \]
\[ \Delta = \log\frac{P(spam)}{P(ham)} + \sum_i \log\frac{P(w_i\mid spam)}{P(w_i\mid ham)} \]

If \(\Delta > 0\), spam wins. If \(\Delta < 0\), ham wins.

Suppose the training data is imbalanced:

\[ P(ham)=0.95 \qquad P(spam)=0.05 \]

The spam class starts with a disadvantage:

\[ \log\frac{P(spam)}{P(ham)} = \log\frac{0.05}{0.95} \approx -2.94 \]

Now consider this new message:

free lottery prize claim

Assume these words provide the following spam evidence:

WordSpam evidence \(\log \frac{P(w\mid spam)}{P(w\mid ham)}\)
free1.8
lottery2.2
prize1.7
claim1.5

Total word evidence:

\[ 1.8+2.2+1.7+1.5=7.2 \]

Final spam-vs-ham evidence:

\[ \Delta = -2.94 + 7.2 = 4.26 \]

Since \(4.26 > 0\), the model predicts spam even though spam is the minority class.

Class imbalance changes the starting point, not the final rule 0 prior penalty: -2.94 word evidence: +7.2 +4.26 spam wins Ham prior says: most emails are ham. Spam words say: this particular email looks very spam-like.

The prior is only the starting belief. The message words can overturn it.

For imbalanced spam datasets, do not judge the model using accuracy alone. Track spam precision, spam recall, and \(F_1\), then choose a threshold based on the business cost of false alarms vs missed spam.

Word evidence in log-space

Interpretability question

Which word in a message contributes most toward spam?

Compare a word’s log likelihood under spam and ham.

\[ \text{spam evidence}(w)=\log P(w\mid spam)-\log P(w\mid ham) \]
ValueMeaning
PositiveThe word pushes toward spam.
NegativeThe word pushes toward ham.
Near zeroThe word is not strongly class-specific.

Unknown words at prediction time

New word question

If the model never saw the word “crypto” during training, what should happen during prediction?

A word outside the training vocabulary has no learned count. Common practical choices are:

ApproachMeaningWhen it is useful
Ignore unknown wordThe word contributes no class evidence.Common with fixed vocabularies such as CountVectorizer.
Use an UNK tokenRare or unseen words map to one shared unknown token.Useful when unknown-word behavior itself carries information.
Use character n-gramsRepresent pieces like prefixes, suffixes, or character windows.Useful for spelling variants, names, and noisy text.
Laplace smoothing protects words inside the learned vocabulary. It does not automatically create probabilities for completely unknown words unless the implementation includes an unknown-token strategy.

Why Naive Bayes is linear in log-space

Boundary question

If each word adds or subtracts evidence, what shape does the final decision take?

For two classes, compare the spam log score and ham log score. The difference can be written as a weighted sum of word counts.

\[ score(spam)-score(ham)=\log\frac{P(spam)}{P(ham)}+\sum_{j=1}^{V}x_j\log\frac{P(w_j\mid spam)}{P(w_j\mid ham)} \]

Here \(x_j\) is the count of word \(j\). This looks like a linear model on bag-of-words features: each word has a fixed evidence weight, and the document adds them up.

This is why Naive Bayes can be very fast: after training, prediction is mostly addition of stored log-evidence weights.

Full training algorithm

Model storage question

After training, what does a Naive Bayes model actually store?

  1. Build the vocabulary from training documents.
  2. Count documents per class to estimate \(P(c)\).
  3. Count words inside each class.
  4. Apply Laplace smoothing to estimate \(P(w\mid c)\).
  5. Store \(\log P(c)\) and \(\log P(w\mid c)\).
model = {
  classes,
  vocabulary,
  class_log_priors,
  word_log_likelihoods
}

Full prediction algorithm

Prediction question

How does the model score a new sentence step by step?

  1. Tokenize the new document.
  2. For each class, start with \(\log P(c)\).
  3. Add \(\log P(w\mid c)\) for every word in the document.
  4. Return the class with the largest log score.
\[ \hat{c}=\arg\max_c\left[\log P(c)+\sum_{i=1}^{n}\log P(w_i\mid c)\right] \]
Previous: Text NBNext: Variants