Instructor view · 56-minute route

Your live-demo runbook.

A structured path that alternates explanation, audience prediction, mathematical derivation, and hands-on experimentation.

Session promise

“By the end, you will not just calculate metrics—you will know which one to choose and why.”

First principlesBusiness contextVisual reasoningLive controlsFrequent checks

Facilitation rule

Predict → demonstrate → explain → check.

Ask learners to commit to an answer before revealing a calculation. This turns every visual into an experiment.

Minute-by-minute route

0–4 min · Hook

Open with the 99% model

Show the overview only. Poll: “Would you deploy it?” Accept both answers, then ask what information is missing.

4–10 min · Motivation

Reveal the accuracy trap

Use the 1,000-transaction example. Establish class imbalance and the need to distinguish mistake types.

10–19 min · Foundation

Construct the confusion matrix

Define “positive” first. Let the panel name one scenario, then adjust the calculator counts together.

19–30 min · Core math

Derive precision, recall, and F1

Start with each metric’s natural-language question. Ask learners to identify the denominator before showing the formula.

30–41 min · Hands-on experiment

Move the threshold and build ROC intuition

Invite a learner to choose a threshold for fraud. Compare operational consequences, then sweep all thresholds conceptually to form ROC.

41–48 min · Rare positives

Contrast ROC-AUC with PR-AUC

Move the operating point on the PR curve and anchor its random baseline to positive-class prevalence.

48–54 min · Regression bridge

Derive R² from a mean baseline

Move the fit slider. Explain 1, 0, and negative values. Avoid claiming R² is “percent accuracy.”

54–56 min · Retrieval recap

Return to the opening model

Ask: “What five questions would you ask before deployment?” Then take questions.

Five-question closing recap

1

How balanced are the classes?

Accuracy needs context.

2

What is the positive class?

The confusion matrix depends on it.

3

Which mistake costs more?

This guides precision versus recall.

4

What threshold can operations support?

AUC does not choose it for us.

5

What baseline did we beat?

R² compares against the mean.

If you have 45 minutes

  • Accuracy trap: 4 min
  • Confusion matrix: 7 min
  • Precision/recall/F1: 10 min
  • Threshold/ROC-AUC: 9 min
  • PR-AUC: 6 min
  • R²: 5 min
  • Recap: 4 min

If you have 60 minutes

  • Use the full 56-minute route.
  • Add audience calculations on the matrix.
  • Compare ROC and PR baselines explicitly.
  • Use remaining time for questions.

Likely panel questions

QuestionCompact answer
Why harmonic mean for F1?It penalizes imbalance: one very low component keeps F1 low.
Can accuracy ever be enough?Yes, when classes and error costs are reasonably balanced—but still inspect the confusion matrix.
Why can ROC-AUC mislead on rare positives?FPR divides by many negatives, so a seemingly small rate can still create many false alarms. Check precision and PR-AUC.
What is the random PR baseline?The positive-class prevalence—not 0.5. With 1% positives, baseline precision is 1%.
Are PR-AUC and Average Precision identical?Not always. They summarize the same curve using different interpolation or weighting conventions.
Does threshold affect ROC-AUC?A single chosen threshold does not change AUC; the ROC curve is built by sweeping all thresholds.
Is R² comparable across every dataset?Not safely. It depends on target variation and evaluation data; use it with error metrics and domain context.

Delivery checklist

Before the panel

  • Open all pages once so assets are cached.
  • Run the live notebook once from top to bottom.
  • Test the interactive sliders.
  • Use 110–125% browser zoom.
  • Keep a five-minute buffer.

During the session

  • Define the positive class aloud.
  • Ask before revealing.
  • Interpret every formula in words.
  • Connect each metric to an action.
Final takeaway: model evaluation is not about finding the biggest number. It is about measuring the mistakes that matter for the decision we need to make.
← R²Return to start ↺