Open with the 99% model
Show the overview only. Poll: “Would you deploy it?” Accept both answers, then ask what information is missing.
A structured path that alternates explanation, audience prediction, mathematical derivation, and hands-on experimentation.
“By the end, you will not just calculate metrics—you will know which one to choose and why.”
Predict → demonstrate → explain → check.
Ask learners to commit to an answer before revealing a calculation. This turns every visual into an experiment.
Show the overview only. Poll: “Would you deploy it?” Accept both answers, then ask what information is missing.
Use the 1,000-transaction example. Establish class imbalance and the need to distinguish mistake types.
Define “positive” first. Let the panel name one scenario, then adjust the calculator counts together.
Start with each metric’s natural-language question. Ask learners to identify the denominator before showing the formula.
Invite a learner to choose a threshold for fraud. Compare operational consequences, then sweep all thresholds conceptually to form ROC.
Move the operating point on the PR curve and anchor its random baseline to positive-class prevalence.
Move the fit slider. Explain 1, 0, and negative values. Avoid claiming R² is “percent accuracy.”
Ask: “What five questions would you ask before deployment?” Then take questions.
Accuracy needs context.
The confusion matrix depends on it.
This guides precision versus recall.
AUC does not choose it for us.
R² compares against the mean.
| Question | Compact answer |
|---|---|
| Why harmonic mean for F1? | It penalizes imbalance: one very low component keeps F1 low. |
| Can accuracy ever be enough? | Yes, when classes and error costs are reasonably balanced—but still inspect the confusion matrix. |
| Why can ROC-AUC mislead on rare positives? | FPR divides by many negatives, so a seemingly small rate can still create many false alarms. Check precision and PR-AUC. |
| What is the random PR baseline? | The positive-class prevalence—not 0.5. With 1% positives, baseline precision is 1%. |
| Are PR-AUC and Average Precision identical? | Not always. They summarize the same curve using different interpolation or weighting conventions. |
| Does threshold affect ROC-AUC? | A single chosen threshold does not change AUC; the ROC curve is built by sweeping all thresholds. |
| Is R² comparable across every dataset? | Not safely. It depends on target variation and evaluation data; use it with error metrics and domain context. |