When accuracy lies.
How do we know whether a machine-learning model is actually good? We will move from raw predictions to business-aware evaluation—one mistake at a time.
A fraud model reports:
accuracy
Would you deploy it?
The one question behind every metric
“What kind of mistake can we afford?”
Metrics are not just formulas. Each metric expresses a preference about which errors matter. By the end, learners should be able to select—not merely calculate—the right metric.
Learning journey
Break the accuracy instinct
Use a 1%-fraud puzzle to discover why “mostly correct” can still be useless.
2Name every outcome
Build the confusion matrix from real decisions and calculate accuracy interactively.
3Choose the costly mistake
Derive precision, recall, and F1 directly from questions—not memorized formulas.
4Move the decision threshold
Watch every metric change, then connect thresholds to ROC and AUC.
5Focus on rare positives
Use the precision–recall curve and PR-AUC when alert quality matters.
6Switch to regression
Understand R² as improvement over the simplest baseline: predicting the mean.
7Run the live session
A minute-by-minute instructor route, audience prompts, recap, and backup plan.
Learning outcomes
By the end, learners will be able to:
- explain why accuracy fails on imbalanced data;
- construct and interpret a confusion matrix;
- calculate precision, recall, and F1;
- describe how a threshold changes model behaviour;
- interpret ROC-AUC as ranking quality;
- choose PR-AUC for rare positive classes;
- explain R² relative to a mean baseline.
Opening audience poll
“A fraud model is 99% accurate. Good model or bad model?”
Ask for a show of hands. Do not resolve it immediately—let the first page reveal the missing information.