A probability is not yet a decision.
The model produces a score. We choose a threshold that turns the score into an action—and that choice changes every cell in the confusion matrix.
Hands-on: move the threshold
Twenty transactions are sorted by fraud score. Red tiles are actual fraud. An orange outline means the model predicts fraud.
Each tile shows the model score × 100. Lower the threshold to flag more cases.
Run the experiment
Lower the threshold. What happens to recall? What price do we usually pay?
Recall generally rises because fewer positives are missed. False positives usually rise too, which can reduce precision.
From one threshold to every threshold
A confusion matrix describes one operating point. An ROC curve repeats the experiment across all thresholds.
- Start with a threshold so high that nothing is positive.
- Lower it one scored case at a time.
- Plot FPR on the x-axis and TPR on the y-axis.
Better curves bow toward the top-left: catch positives without many false alarms.
What AUC means
Area Under the ROC Curve summarizes ranking quality across all thresholds.
An intuitive interpretation: ROC-AUC is the probability that a randomly chosen positive receives a higher score than a randomly chosen negative.
What AUC does not mean
- It does not choose a deployment threshold.
- It does not encode the business cost of mistakes.
- It can look optimistic with extreme class imbalance.
Decision question
Two models have the same ROC-AUC. Could one still be better for fraud operations?
Yes. They can behave differently near the threshold the business can actually use. Compare precision, recall, capacity, calibration, and error cost at the intended operating point.