Act 3 · Core metrics · 12 minutes

Two denominators. Two different questions.

Precision and recall use the same numerator—true positives—but look at two different groups.

Start from predictions

Precision

“When the model raises an alarm, how often is it right?”

\[\text{Precision}=\frac{TP}{TP+FP}\]

The denominator is everything predicted positive. False alarms reduce precision.

Prioritize precision when acting on a positive prediction is expensive: manual fraud reviews, blocking email, or invasive follow-up tests.
Start from reality

Recall

“Of all real positive cases, how many did the model catch?”

\[\text{Recall}=\frac{TP}{TP+FN}\]

The denominator is everything actually positive. Missed cases reduce recall.

Prioritize recall when missing a positive is dangerous: disease screening, fraud detection, or safety failures.

A memory aid that follows the denominator

Predicted fraud
Predicted legitimate
Actual fraud
TP = 8caught
FN = 2missed
Actual legitimate
FP = 12false alarms
TN = 78cleared
\[\text{Precision}=\frac{8}{8+12}=40\%\]
\[\text{Recall}=\frac{8}{8+2}=80\%\]

Interpret before calculating

Is this model better at catching fraud, or at making trustworthy fraud alerts?

It catches 80% of fraud, but only 40% of its alerts are genuine. It has stronger recall than precision.

Why not simply average them?

The arithmetic mean can remain high even when one metric is weak. F1 uses the harmonic mean, which is pulled toward the smaller value.

\[F1=2\cdot\frac{\text{Precision}\cdot\text{Recall}}{\text{Precision}+\text{Recall}}\]
\[F1=2\cdot\frac{0.40\times0.80}{0.40+0.80}=0.53\]
PrecisionRecallF140%80%53%

F1 becomes high only when both precision and recall are reasonably high.

Metric choice is a product decision

ScenarioCostlier mistakeMetric to emphasize
Medical screeningMissing a disease (FN)Recall
Spam auto-deleteDeleting a real email (FP)Precision
Fraud investigation queueBoth missed fraud and wasted reviewsF1 or cost-based metric
There is no universally best metric. The right metric reflects class balance, error cost, and what action follows a prediction.
← Confusion matrixNext: Thresholds & ROC-AUC →