Act 5 · Rare-event evaluation · 8 minutes

PR-AUC focuses on the class we care about.

When positives are rare, we want to know two things: how many positives we catch, and whether our positive alerts are trustworthy.

What the curve plots

\[x=\text{Recall}=\frac{TP}{TP+FN}\]
\[y=\text{Precision}=\frac{TP}{TP+FP}\]

Each threshold creates one point. Sweeping from a high threshold to a low threshold traces the precision–recall curve.

Why it helps with rare positives

Both axes focus on positive-class performance. True negatives—which can dominate a highly imbalanced dataset—do not appear in either formula.

Question answered: As we try to catch more positives, how trustworthy do our alerts remain?

Hands-on: move along the PR curve

Change the fraud threshold. The orange point shows the current precision and recall. The complete blue path summarizes all thresholds.

0.50
RecallPrecisionrandom baseline

Toward the right means more positives caught. Toward the top means more trustworthy alerts.

—Precision
—Recall
—Alerts

Experiment

Lower the threshold until recall rises. What happens to alert volume and precision?

You catch more positives, but usually create more false alarms. The PR curve makes this trade-off visible.

Reading the score

PR-AUC summarizes the area under the precision–recall curve. Higher is better, but its baseline depends on how common the positive class is.

\[\text{random baseline precision}=\frac{\text{positives}}{\text{all examples}}\]
—Positive rate
—Average precision
1.00Perfect

Baseline intuition

If fraud is 1% of all transactions, random alerts have about 1% precision—not 50%.

That is why a PR score must be interpreted relative to prevalence and compared on the same evaluation dataset.

A PR-AUC of 0.40 may be excellent when prevalence is 0.01, but disappointing when prevalence is 0.35.

ROC-AUC versus PR-AUC

ROC-AUCPR-AUC
AxesTPR versus FPRPrecision versus recall
FocusRanking positives above negativesQuality and coverage of positive predictions
Best useOverall ranking, moderate class balanceRare positive class and alert quality
Random baseline0.50Positive-class prevalence
Main cautionCan look optimistic under severe imbalanceChanges when class prevalence changes

PR-AUC and Average Precision

They are closely related summaries, but not always numerically identical. Trapezoidal PR-AUC linearly interpolates between curve points. Average Precision (AP) uses a step-weighted summary based on recall gains.

Many libraries and articles use the terms loosely, so check the exact function being reported.

Python implementation

from sklearn.metrics import (
    precision_recall_curve,
    average_precision_score
)

precision, recall, thresholds = precision_recall_curve(
    y_true, y_score
)

ap = average_precision_score(y_true, y_score)

Check for understanding

For a disease affecting 0.5% of people, should a random classifier’s PR baseline be 0.5 or 0.005?

0.005. Baseline precision equals prevalence: 0.5% = 0.005. ROC’s random baseline is 0.5, but PR’s is not.

← ROC-AUCNext: R² for regression →