Bayes rule turns evidence into belief.
Naive Bayes starts with a prior belief about classes, then updates that belief after seeing features or words.
Prior probability
Before seeing the message
If 80 out of 100 messages are normal and 20 are spam, what should be the model’s first guess?
The prior is the model’s starting belief before looking at the message.
| Class | Count | Prior |
|---|---|---|
| ham | 80 | \(P(ham)=0.80\) |
| spam | 20 | \(P(spam)=0.20\) |
Prior is the base rate. It matters before we look at words.
Likelihood
Evidence question
If we already know a message is spam, how likely is the word “free”?
A likelihood measures how expected the evidence is under a class.
This does not mean “probability of spam given free”. It means “probability of seeing the word free if the class is spam”.
| Quantity | Question it answers |
|---|---|
| \(P(spam)\) | How common is spam before seeing this message? |
| \(P(free \mid spam)\) | If this is spam, how expected is the word free? |
| \(P(spam \mid free)\) | After seeing free, how likely is spam? |
Bayes rule
Turn it around
Training gives us \(P(word\mid class)\), but prediction needs \(P(class\mid word)\). What connects them?
| Part | Name | Meaning |
|---|---|---|
| \(P(c\mid d)\) | Posterior | Belief about class after seeing document. |
| \(P(c)\) | Prior | Belief about class before seeing document. |
| \(P(d\mid c)\) | Likelihood | How expected the document is if class is known. |
| \(P(d)\) | Evidence | Overall probability of this document. |
Why the denominator can be ignored
Comparison question
When comparing spam vs ham for the same message, does \(P(d)\) change?
No. The message is fixed, so \(P(d)\) is the same for every class. For choosing the winning class, we can compare only the numerator.
Generative intuition
Reverse story
Could we imagine the class being chosen first, and then the words being generated from that class?
Naive Bayes is often explained as a generative classifier. It imagines this process:
Prediction reverses the story: after observing the document, the model asks which class most likely generated it.
Generative vs discriminative models
Modeling question
Should a model directly learn the class boundary, or should it learn how each class could have produced the data?
This is the key difference between discriminative and generative models.
| Model type | What it learns | Question it asks | Examples |
|---|---|---|---|
| Discriminative | \(P(y\mid x)\) | Given this input, which class is it? | Logistic Regression, SVM, many neural classifiers |
| Generative | \(P(x\mid y)\) and \(P(y)\) | If this class were true, how likely is this input? | Naive Bayes, Gaussian Mixture Models, Hidden Markov Models |
How Naive Bayes is generative
Story question
Imagine this email was spam. Are these words likely? Now imagine it was ham. Are these words likely?
Naive Bayes imagines that the class is chosen first, and then the words are generated from that class.
- Choose a class using \(P(c)\).
- Generate words using \(P(w\mid c)\).
- For a new message, choose the class that makes the words most expected.
For text, the naive assumption turns \(P(x\mid c)\) into a product of word likelihoods.
Naive Bayes learns a simple class-wise word-generation model.
Worked comparison example
Evidence question
If ham is more common, can spam words still overcome the ham prior?
Suppose:
And the model has learned these word likelihoods:
| Word | \(P(w\mid spam)\) | \(P(w\mid ham)\) |
|---|---|---|
| free | 0.20 | 0.01 |
| lottery | 0.10 | 0.005 |
| prize | 0.15 | 0.01 |
For the message free lottery prize:
Even though ham has the larger prior, the words are much more likely under spam. So the predicted class is:
What can we do with a generative model?
Beyond prediction
If the model learns \(P(w\mid spam)\) and \(P(w\mid ham)\), what else can we inspect?
| Use | What it means in Naive Bayes |
|---|---|
| Inspect class behavior | Find words with high \(P(w\mid spam)\) or high \(P(w\mid ham)\). |
| Explain predictions | Show which words contributed evidence toward each class. |
| Generate examples conceptually | Sample likely words from a class to imagine spam-like or ham-like messages. |
| Use priors explicitly | Adjust \(P(c)\) if real-world class frequencies or business needs differ. |
| Handle missing features naturally | Score the observed features and ignore missing ones when appropriate. |