PR-AUC focuses on the class we care about.
When positives are rare, we want to know two things: how many positives we catch, and whether our positive alerts are trustworthy.
What the curve plots
Each threshold creates one point. Sweeping from a high threshold to a low threshold traces the precision–recall curve.
Why it helps with rare positives
Both axes focus on positive-class performance. True negatives—which can dominate a highly imbalanced dataset—do not appear in either formula.
Hands-on: move along the PR curve
Change the fraud threshold. The orange point shows the current precision and recall. The complete blue path summarizes all thresholds.
Toward the right means more positives caught. Toward the top means more trustworthy alerts.
Experiment
Lower the threshold until recall rises. What happens to alert volume and precision?
You catch more positives, but usually create more false alarms. The PR curve makes this trade-off visible.
Reading the score
PR-AUC summarizes the area under the precision–recall curve. Higher is better, but its baseline depends on how common the positive class is.
Baseline intuition
If fraud is 1% of all transactions, random alerts have about 1% precision—not 50%.
That is why a PR score must be interpreted relative to prevalence and compared on the same evaluation dataset.
ROC-AUC versus PR-AUC
| ROC-AUC | PR-AUC | |
|---|---|---|
| Axes | TPR versus FPR | Precision versus recall |
| Focus | Ranking positives above negatives | Quality and coverage of positive predictions |
| Best use | Overall ranking, moderate class balance | Rare positive class and alert quality |
| Random baseline | 0.50 | Positive-class prevalence |
| Main caution | Can look optimistic under severe imbalance | Changes when class prevalence changes |
PR-AUC and Average Precision
They are closely related summaries, but not always numerically identical. Trapezoidal PR-AUC linearly interpolates between curve points. Average Precision (AP) uses a step-weighted summary based on recall gains.
Many libraries and articles use the terms loosely, so check the exact function being reported.
Python implementation
from sklearn.metrics import (
precision_recall_curve,
average_precision_score
)
precision, recall, thresholds = precision_recall_curve(
y_true, y_score
)
ap = average_precision_score(y_true, y_score)Check for understanding
For a disease affecting 0.5% of people, should a random classifier’s PR baseline be 0.5 or 0.005?
0.005. Baseline precision equals prevalence: 0.5% = 0.005. ROC’s random baseline is 0.5, but PR’s is not.