Two denominators. Two different questions.
Precision and recall use the same numerator—true positives—but look at two different groups.
Precision
“When the model raises an alarm, how often is it right?”
The denominator is everything predicted positive. False alarms reduce precision.
Recall
“Of all real positive cases, how many did the model catch?”
The denominator is everything actually positive. Missed cases reduce recall.
A memory aid that follows the denominator
Interpret before calculating
Is this model better at catching fraud, or at making trustworthy fraud alerts?
It catches 80% of fraud, but only 40% of its alerts are genuine. It has stronger recall than precision.
Why not simply average them?
The arithmetic mean can remain high even when one metric is weak. F1 uses the harmonic mean, which is pulled toward the smaller value.
F1 becomes high only when both precision and recall are reasonably high.
Metric choice is a product decision
| Scenario | Costlier mistake | Metric to emphasize |
|---|---|---|
| Medical screening | Missing a disease (FN) | Recall |
| Spam auto-delete | Deleting a real email (FP) | Precision |
| Fraud investigation queue | Both missed fraud and wasted reviews | F1 or cost-based metric |