Text Naive Bayes learns from word counts.
The simplest strong version for text is Multinomial Naive Bayes: represent each document as word counts, then learn word probabilities inside each class.
Tokenization
First representation question
What should a model count from the sentence “Free prize claim now”?
Tokenization breaks text into useful pieces, usually words or subwords. For a first class example, simple lowercase words are enough.
The model cannot use raw text directly. It needs tokens and counts.
Vocabulary and bag-of-words
Counting question
If word order is ignored, what information remains?
The vocabulary is the set of words learned from training data. Bag-of-words converts each document into a count vector over that vocabulary.
Here \(x_j\) is the count of vocabulary word \(j\), and \(V\) is vocabulary size.
| Message | free | prize | meeting | project | claim |
|---|---|---|---|---|---|
| free prize claim | 1 | 1 | 0 | 0 | 1 |
| project meeting today | 0 | 0 | 1 | 1 | 0 |
| free free prize | 2 | 1 | 0 | 0 | 0 |
Class-wise word counts
Evidence question
Which words should push a prediction toward spam, and which words should push it toward ham?
Naive Bayes counts words separately inside each class.
| Word | Count in ham | Count in spam | Intuition |
|---|---|---|---|
| free | 0 | 4 | Strong spam evidence |
| claim | 0 | 3 | Strong spam evidence |
| meeting | 4 | 0 | Strong ham evidence |
| project | 3 | 0 | Strong ham evidence |
| today | 2 | 1 | Weak evidence |
A word’s class-wise count difference gives an early sense of evidence.
Likelihood from counts
Probability question
If spam messages contain 30 total words and the word “free” appears 4 times, what is \(P(free\mid spam)\)?
Example:
The naive assumption
Independence question
If a message contains “free”, does that make “prize” more likely too?
In real language, yes. Words are related. But Naive Bayes simplifies the problem by pretending that once the class is known, each word gives separate evidence.
This assumption is usually false, but useful. It avoids needing a probability for every possible sentence.
Prediction formula
Decision question
For “free prize today”, should the model multiply word evidence under ham and spam, then compare?
For the same message, the model computes one score for each class and picks the larger score.
| Class | Score shape | Interpretation |
|---|---|---|
| ham | \(P(ham)P(free\mid ham)P(prize\mid ham)P(today\mid ham)\) | How well ham explains the words. |
| spam | \(P(spam)P(free\mid spam)P(prize\mid spam)P(today\mid spam)\) | How well spam explains the words. |