Naive Bayes: classify by asking which class best explains the evidence.
Naive Bayes is a fast, interpretable probabilistic classifier. It is especially useful for text classification, where a document becomes evidence made from words.
Why should we learn it?
Opening question
When you see a message saying “free cash prize”, why do you immediately suspect spam?
You are mentally using word evidence. Naive Bayes formalizes this intuition using probability.
Instead of learning slopes like linear regression, Naive Bayes learns probability tables: how common each class is, and how common each word or feature is inside each class.
A useful mental model: the model asks, “If this message was spam, how expected are these words? If it was ham, how expected are these words?”
Naive Bayes compares class-wise evidence, not geometric distance.
Session story
We begin with intuition, then build the exact formula, then handle practical details like smoothing, log-space computation, variants, and evaluation.
Roadmap
Bayes intuition
Prior, likelihood, posterior, and why the denominator can be ignored for classification.
Text Naive Bayes
Tokenization, vocabulary, bag-of-words, class-wise word counts, and the naive assumption.
Smoothing and logs
Zero probabilities, Laplace smoothing, log probabilities, and the full training/prediction algorithm.
Variants and evaluation
Multinomial, Bernoulli, Gaussian NB, text preprocessing choices, metrics, strengths, and limitations.
Where Naive Bayes appears
| Problem | Why Naive Bayes fits |
|---|---|
| Spam detection | Words like “free”, “claim”, and “winner” carry class evidence. |
| Sentiment classification | Words and phrases provide evidence for positive or negative sentiment. |
| Topic classification | Words such as “orbit”, “team”, or “market” point toward topics. |
| Language identification | Character or word patterns differ strongly across languages. |
| Medical screening baseline | Symptoms or test signals can be combined probabilistically, with care. |
Reference idea
This material is strengthened using the Naive Bayes and text classification treatment from Jurafsky and Martin, Speech and Language Processing, Chapter B, especially the generative-classifier view, bag-of-words assumption, smoothing, log-space scoring, and evaluation metrics.