The 99% accuracy trap.
Before learning formulas, let us discover why evaluation needs more than “how often was the model correct?”
Imagine 1,000 transactions
Only 10 are fraudulent. The remaining 990 are legitimate.
Think–pair–share
A lazy model predicts “legitimate” for every transaction. What accuracy will it receive?
990 ÷ 1,000 = 99% accuracy
But it catches zero of the 10 frauds. Its recall for fraud is 0%. High accuracy; zero usefulness.
The class we care about can be visually—and statistically—drowned out by the majority.
First principle: a metric compresses reality
When accuracy is useful
Classes are reasonably balanced and false positives and false negatives have similar costs.
When accuracy is dangerous
One class is rare, or one kind of mistake is much more costly than another.
Bridge question
If total correctness hides the important mistakes, what should we count instead?
We need to separate correct and incorrect predictions by the actual class. That gives us the confusion matrix.