Part 2

Text Naive Bayes learns from word counts.

The simplest strong version for text is Multinomial Naive Bayes: represent each document as word counts, then learn word probabilities inside each class.

Tokenization

First representation question

What should a model count from the sentence “Free prize claim now”?

Tokenization breaks text into useful pieces, usually words or subwords. For a first class example, simple lowercase words are enough.

\[ \text{"Free prize claim now"}\rightarrow[free, prize, claim, now] \]
freeprizeclaimnow
Free prize claim now free prize claim now

The model cannot use raw text directly. It needs tokens and counts.

Vocabulary and bag-of-words

Counting question

If word order is ignored, what information remains?

The vocabulary is the set of words learned from training data. Bag-of-words converts each document into a count vector over that vocabulary.

\[ d=[x_1,x_2,\ldots,x_V] \]

Here \(x_j\) is the count of vocabulary word \(j\), and \(V\) is vocabulary size.

Messagefreeprizemeetingprojectclaim
free prize claim11001
project meeting today00110
free free prize21000
Bag-of-words is simple and powerful, but it loses word order. “not good” and “good” need extra preprocessing if negation matters.

Class-wise word counts

Evidence question

Which words should push a prediction toward spam, and which words should push it toward ham?

Naive Bayes counts words separately inside each class.

WordCount in hamCount in spamIntuition
free04Strong spam evidence
claim03Strong spam evidence
meeting40Strong ham evidence
project30Strong ham evidence
today21Weak evidence
Spam count minus ham count ham-like spam-like free claim meeting project today

A word’s class-wise count difference gives an early sense of evidence.

Likelihood from counts

Probability question

If spam messages contain 30 total words and the word “free” appears 4 times, what is \(P(free\mid spam)\)?

\[ P(w\mid c)=\frac{\text{count of word }w\text{ in class }c}{\text{total word count in class }c} \]

Example:

\[ P(free\mid spam)=\frac{4}{30}=0.133 \]

The naive assumption

Independence question

If a message contains “free”, does that make “prize” more likely too?

In real language, yes. Words are related. But Naive Bayes simplifies the problem by pretending that once the class is known, each word gives separate evidence.

\[ P(w_1,w_2,\ldots,w_n\mid c)\approx\prod_{i=1}^{n}P(w_i\mid c) \]

This assumption is usually false, but useful. It avoids needing a probability for every possible sentence.

The algorithm is “naive” because of this simplifying independence assumption, not because it is weak.

Prediction formula

Decision question

For “free prize today”, should the model multiply word evidence under ham and spam, then compare?

\[ \hat{c}=\arg\max_c P(c)\prod_{i=1}^{n}P(w_i\mid c) \]

For the same message, the model computes one score for each class and picks the larger score.

ClassScore shapeInterpretation
ham\(P(ham)P(free\mid ham)P(prize\mid ham)P(today\mid ham)\)How well ham explains the words.
spam\(P(spam)P(free\mid spam)P(prize\mid spam)P(today\mid spam)\)How well spam explains the words.
Previous: BayesNext: Smoothing