A model must beat a sensible guess.
A baseline is a complete forecasting method, not a decorative score. It encodes the simplest plausible story about persistence, repetition or growth and gives every complex model a fair opponent.
Tomorrow question
If today is Sunday, which is a better guess for Monday: Sunday's demand or last Monday's demand? What assumption does each choice make?
Five useful baselines
Mean
Assumes no useful time structure beyond a stable level.
Naive
Assumes the most recently observed level persists.
Seasonal naive
Repeats the matching position from the latest observed season, where \(k=\lfloor(h-1)/m\rfloor\).
Drift
Extends the average change from the first to latest observation.
Recent-window mean
Uses the average of the latest \(w\) observations as the next level.
Terminology: this recent-window mean is often called a moving-average forecast. It is not the MA(\(q\)) error model inside ARIMA.
Calculate before revealing
Fourteen daily observations end on Sunday. Forecast the next Monday using naive, seasonal naive, drift and a three-day mean.
A worked baseline example
| Week | Mon | Tue | Wed | Thu | Fri | Sat | Sun |
|---|---|---|---|---|---|---|---|
| Week 1 | 100 | 110 | 120 | 125 | 140 | 170 | 155 |
| Week 2 | 108 | 116 | 127 | 132 | 148 | 180 | 164 |
Naive
“Monday will look like yesterday.”
Seasonal naive
“Monday will look like last Monday.”
Drift
“The average upward change will continue.”
Three-day mean
“The recent local level will continue.”
Suppose the next Monday is \(y_{15}=116\). Absolute errors are:
| Method | Forecast | Absolute error |
|---|---|---|
| Naive | 164 | 48 |
| Seasonal naive | 108 | 8 |
| Drift | 168.92 | 52.92 |
| Three-day mean | 164 | 48 |
Seasonal naive wins this origin because the day-of-week pattern is stronger than yesterday's level or overall drift. One origin is not enough to select it; repeat the comparison across many historical forecast origins.
Hold out the final two weeks
The shaded region is unseen at the forecast origin. Both methods generate the full 14-day path using only the first 42 days; seasonal naive cycles through the final observed week rather than reading actual values from the holdout.
Interpretation: the last value ignores day-of-week structure. Seasonal naive follows the weekly shape and becomes the benchmark a more complex model must improve upon.
Fairness question
Can we select the best baseline after looking at the final test period and still call it an untouched test set?
Match the baseline to the structure
| Observed behavior | First baseline | Reason |
|---|---|---|
| Stable level, little dependence | Training mean | Tests whether temporal structure adds value. |
| Stable local level | Naive | The newest observation is informative. |
| Strong daily cycle in hourly data | Seasonal naive with \(m=24\) | Compare the same hour yesterday. |
| Strong weekly cycle in hourly data | Seasonal naive with \(m=168\) | Compare the same hour and weekday last week. |
| Persistent trend, weak seasonality | Drift | Projects average historical change. |