Naive Bayes Classifier

Naive Bayes: classify by asking which class best explains the evidence.

Naive Bayes is a fast, interpretable probabilistic classifier. It is especially useful for text classification, where a document becomes evidence made from words.

Why should we learn it?

Opening question

When you see a message saying “free cash prize”, why do you immediately suspect spam?

You are mentally using word evidence. Naive Bayes formalizes this intuition using probability.

Instead of learning slopes like linear regression, Naive Bayes learns probability tables: how common each class is, and how common each word or feature is inside each class.

\[ \text{prediction} = \text{class with strongest prior + evidence} \]

A useful mental model: the model asks, “If this message was spam, how expected are these words? If it was ham, how expected are these words?”

Message evidence free prize claim today ham? meeting, project spam? free, prize, claim The class that makes the words most expected wins.

Naive Bayes compares class-wise evidence, not geometric distance.

Session story

Start with a message Count words Learn probabilities Add log evidence Evaluate decisions

We begin with intuition, then build the exact formula, then handle practical details like smoothing, log-space computation, variants, and evaluation.

Roadmap

Where Naive Bayes appears

ProblemWhy Naive Bayes fits
Spam detectionWords like “free”, “claim”, and “winner” carry class evidence.
Sentiment classificationWords and phrases provide evidence for positive or negative sentiment.
Topic classificationWords such as “orbit”, “team”, or “market” point toward topics.
Language identificationCharacter or word patterns differ strongly across languages.
Medical screening baselineSymptoms or test signals can be combined probabilistically, with care.

Reference idea

This material is strengthened using the Naive Bayes and text classification treatment from Jurafsky and Martin, Speech and Language Processing, Chapter B, especially the generative-classifier view, bag-of-words assumption, smoothing, log-space scoring, and evaluation metrics.

Next: Bayes Intuition