Part 1

Bayes rule turns evidence into belief.

Naive Bayes starts with a prior belief about classes, then updates that belief after seeing features or words.

Prior probability

Before seeing the message

If 80 out of 100 messages are normal and 20 are spam, what should be the model’s first guess?

The prior is the model’s starting belief before looking at the message.

\[ P(c)=\frac{\text{number of training examples in class }c}{\text{total number of training examples}} \]
ClassCountPrior
ham80\(P(ham)=0.80\)
spam20\(P(spam)=0.20\)
Prior probability ham spam 0.80 0.20

Prior is the base rate. It matters before we look at words.

Likelihood

Evidence question

If we already know a message is spam, how likely is the word “free”?

A likelihood measures how expected the evidence is under a class.

\[ P(\text{free}\mid spam) \]

This does not mean “probability of spam given free”. It means “probability of seeing the word free if the class is spam”.

QuantityQuestion it answers
\(P(spam)\)How common is spam before seeing this message?
\(P(free \mid spam)\)If this is spam, how expected is the word free?
\(P(spam \mid free)\)After seeing free, how likely is spam?

Bayes rule

Turn it around

Training gives us \(P(word\mid class)\), but prediction needs \(P(class\mid word)\). What connects them?

\[ P(c\mid d)=\frac{P(c)P(d\mid c)}{P(d)} \]
PartNameMeaning
\(P(c\mid d)\)PosteriorBelief about class after seeing document.
\(P(c)\)PriorBelief about class before seeing document.
\(P(d\mid c)\)LikelihoodHow expected the document is if class is known.
\(P(d)\)EvidenceOverall probability of this document.

Why the denominator can be ignored

Comparison question

When comparing spam vs ham for the same message, does \(P(d)\) change?

No. The message is fixed, so \(P(d)\) is the same for every class. For choosing the winning class, we can compare only the numerator.

\[ \hat{c}=\arg\max_c P(c\mid d)=\arg\max_c P(c)P(d\mid c) \]
We are not saying \(P(d)\) is unimportant mathematically. We are saying it is the same multiplier for every class during argmax classification.

Generative intuition

Reverse story

Could we imagine the class being chosen first, and then the words being generated from that class?

Naive Bayes is often explained as a generative classifier. It imagines this process:

Choose classPick word 1Pick word 2Pick word 3Observe document

Prediction reverses the story: after observing the document, the model asks which class most likely generated it.

Generative vs discriminative models

Modeling question

Should a model directly learn the class boundary, or should it learn how each class could have produced the data?

This is the key difference between discriminative and generative models.

Model typeWhat it learnsQuestion it asksExamples
Discriminative\(P(y\mid x)\)Given this input, which class is it?Logistic Regression, SVM, many neural classifiers
Generative\(P(x\mid y)\) and \(P(y)\)If this class were true, how likely is this input?Naive Bayes, Gaussian Mixture Models, Hidden Markov Models
Short version: discriminative models learn the boundary. Generative models learn the class-wise data story.

How Naive Bayes is generative

Story question

Imagine this email was spam. Are these words likely? Now imagine it was ham. Are these words likely?

Naive Bayes imagines that the class is chosen first, and then the words are generated from that class.

  1. Choose a class using \(P(c)\).
  2. Generate words using \(P(w\mid c)\).
  3. For a new message, choose the class that makes the words most expected.
\[ \hat{c}=\arg\max_c P(c)P(x\mid c) \]

For text, the naive assumption turns \(P(x\mid c)\) into a product of word likelihoods.

\[ P(x\mid c)\approx\prod_{i=1}^{n}P(w_i\mid c) \]
Generative story Choose class Generate words free lottery prize Prediction reverses the story: which class most likely produced these words?

Naive Bayes learns a simple class-wise word-generation model.

Worked comparison example

Evidence question

If ham is more common, can spam words still overcome the ham prior?

Suppose:

\[ P(spam)=0.20 \qquad P(ham)=0.80 \]

And the model has learned these word likelihoods:

Word\(P(w\mid spam)\)\(P(w\mid ham)\)
free0.200.01
lottery0.100.005
prize0.150.01

For the message free lottery prize:

\[ score(spam)=0.20\times0.20\times0.10\times0.15=0.0006 \]
\[ score(ham)=0.80\times0.01\times0.005\times0.01=0.0000004 \]

Even though ham has the larger prior, the words are much more likely under spam. So the predicted class is:

\[ \hat{c}=spam \]

What can we do with a generative model?

Beyond prediction

If the model learns \(P(w\mid spam)\) and \(P(w\mid ham)\), what else can we inspect?

UseWhat it means in Naive Bayes
Inspect class behaviorFind words with high \(P(w\mid spam)\) or high \(P(w\mid ham)\).
Explain predictionsShow which words contributed evidence toward each class.
Generate examples conceptuallySample likely words from a class to imagine spam-like or ham-like messages.
Use priors explicitlyAdjust \(P(c)\) if real-world class frequencies or business needs differ.
Handle missing features naturallyScore the observed features and ignore missing ones when appropriate.
Naive Bayes is generative, but basic text Naive Bayes is not a modern text generator. It can describe class-wise word probabilities; it does not generate fluent language like an LLM.
Previous: OverviewNext: Text NB