Act 1 · Motivation · 6 minutes

The 99% accuracy trap.

Before learning formulas, let us discover why evaluation needs more than “how often was the model correct?”

Imagine 1,000 transactions

Only 10 are fraudulent. The remaining 990 are legitimate.

1,000transactions
10fraud
990legitimate

Think–pair–share

A lazy model predicts “legitimate” for every transaction. What accuracy will it receive?

990 ÷ 1,000 = 99% accuracy

But it catches zero of the 10 frauds. Its recall for fraud is 0%. High accuracy; zero usefulness.

Grey = legitimate (990) · Red = fraud (10)

The class we care about can be visually—and statistically—drowned out by the majority.

First principle: a metric compresses reality

1,000 decisionsDifferent mistakesOne summary numberPossible blind spot
Accuracy answers: “What fraction of all predictions were correct?” It does not tell us which cases were missed.
Accuracy = correct predictions ÷ all predictions

When accuracy is useful

Classes are reasonably balanced and false positives and false negatives have similar costs.

When accuracy is dangerous

One class is rare, or one kind of mistake is much more costly than another.

Bridge question

If total correctness hides the important mistakes, what should we count instead?

We need to separate correct and incorrect predictions by the actual class. That gives us the confusion matrix.

← OverviewNext: Name every outcome →