Model a stable mechanism, not a moving target.
Stationarity asks whether the statistical behavior observed in one period remains a useful description of another period. It is the bridge between “this happened in the past” and “this relationship may repeat in the future.”
Window-shift question
Imagine hiding the date labels and showing three different months from the series. Would their level, spread, and short-term movement look as though they came from the same process?
What is stationarity?
A time series is stationary when its probability-generating mechanism does not change as time moves forward. The values may rise and fall, but the rules governing those movements remain stable.
Classroom intuition: slide a window along the timeline. A stationary series should show broadly similar center, variability, and dependence in each window. It does not need to repeat the exact same values.
Stationary-looking
\(98,103,100,96,102,101,99,\ldots\)
Values fluctuate around approximately 100 with a similar spread throughout.
Non-stationary trend
\(100,108,116,124,132,140,\ldots\)
The typical level increases, so an early-window mean no longer describes a later window.
The first series changes value without changing its behavior. The second changes its mean; the third keeps a similar center but its fluctuations grow. Only the first is stationary-looking.
Meaning question
Does stationary mean “constant,” “without noise,” or “perfectly predictable”?
Weak stationarity, one condition at a time
1. Stable mean
The expected center does not depend on calendar time. Demand cannot average 100 early in the series and 300 later while satisfying this condition.
2. Stable variance
The typical size of fluctuations remains stable. A quiet first year followed by a highly volatile second year violates this condition.
3. Stable memory
Dependence is determined by the distance \(k\), not the date. The one-day relationship should be comparable in January and June.
What this explains: a relationship estimated from historical windows can be reused at a future forecast origin. If the mean, variance, or lag relationship keeps changing, one fixed model is trying to learn several different processes.
Weak versus strict stationarity: strict stationarity requires the complete joint distribution to remain unchanged after shifting time. In practical forecasting, weak stationarity is more commonly discussed because it focuses on mean, variance, and covariance.
Generalization question
A model learns that demand usually returns toward 100. If the business expands and the new normal becomes 200, should that old relationship still be trusted?
Why stationarity matters in forecasting
A forecasting model learns relationships from earlier observations and applies them at a later forecast origin. This transfer from past to future is credible only when the learned mechanism remains reasonably stable.
Core intuition: stationarity does not make the future certain. It makes the past statistically relevant. Without some form of stability, excellent historical fit may describe a regime that no longer exists.
1. Models can generalize
A stable mean, variance, and lag structure allow relationships estimated in training windows to remain useful in validation and deployment.
2. Parameters keep their meaning
An AR coefficient should describe the same type of dependence throughout the series. If dependence changes with time, one fixed coefficient becomes an average of incompatible regimes.
3. ACF and PACF are interpretable
Classical autocorrelation assumes that dependence at lag k is comparable wherever it is measured. Trend and changing seasonality can create correlations that reflect shared time movement rather than useful memory.
4. Uncertainty is more credible
If error variability is stable, historical residuals can inform future prediction intervals. When variance grows, intervals learned from quiet periods become too narrow.
5. Residual checks become meaningful
After a model removes predictable structure, its residuals should behave like a stable, patternless process. Remaining trend, seasonality, or changing spread signals unfinished modeling.
6. Forecast errors reveal change
When a previously stable model develops persistent bias, it can indicate drift, a structural break, or a new operating regime rather than ordinary random error.
A model trained around the wrong level
Suppose historical demand fluctuates around 100 and we estimate:
If yesterday's demand is 120, the one-step forecast is:
The model expects deviations to move back toward 100. But if expansion permanently changes the normal level to 200, that mean-reversion target is obsolete. The same fitted equation will systematically underforecast.
Important qualification: the raw series does not always need to be stationary. Exponential smoothing can model an evolving level, regression can model trend and calendar effects, and ARIMA can difference the series. What must remain stable is the mechanism left for the model to learn after those components are represented.
Examples and non-examples
| Series | Stationary? | Reason |
|---|---|---|
| Noise fluctuating around 100 with stable spread | Possibly | Its level and variance can remain stable even though individual values are unpredictable. |
| Bike demand growing as the service expands | No | The mean changes through time because of trend. |
| Hourly demand with a repeated morning peak | Not in raw form | The expected value depends on the hour. Seasonal adjustment or seasonal differencing may reveal a stable remainder. |
| Sales whose fluctuations grow with sales level | No | The variance changes; a logarithm may stabilize proportional variation. |
| Random walk: \(Y_t=Y_{t-1}+\varepsilon_t\) | No | Every shock permanently changes the level, and \(\operatorname{Var}(Y_t)\) grows with time. |
| First difference of a random walk: \(\Delta Y_t=\varepsilon_t\) | Often yes | Differencing removes the accumulated level and leaves the innovations. |
Seasonality nuance: deterministic seasonality is predictable rather than random, but it still makes the unconditional mean depend on time position. We can model the seasonal state directly or remove it before fitting a model that assumes stationarity.
Diagnosis question
If the first and second half have similar averages, is that enough to declare the series stationary?
How to identify stationarity
- Plot the values in time order. Look for trend, repeating seasonal level, growing spread, sudden breaks, and long periods at different levels. A time plot is more informative than a histogram because it preserves order.
- Compare time windows. Split the series into early, middle, and late windows. Compare their mean and standard deviation. Large systematic differences suggest changing level or variance.
- Plot rolling statistics. A rolling mean exposes local level; a rolling standard deviation exposes local spread. Use them as visual diagnostics, not formal proof, because their appearance depends on window length.
- Inspect the ACF. A trending or unit-root series often has large positive autocorrelations that decay very slowly. Seasonal peaks indicate that the mean depends on calendar position. A stationary series usually loses correlation more quickly, although persistent stationary processes can still decay slowly.
- Use statistical tests. ADF tests a unit-root null; KPSS tests a stationarity null. Their hypotheses differ, so using both can expose uncertainty rather than forcing one test to make the entire decision.
- Repeat after transformation. If you apply a log, first difference, or seasonal difference, replot the result and rerun the diagnostics. The transformed series, not the original one, is what the stationary model will see.
| Diagnostic observation | What it suggests | What it does not prove |
|---|---|---|
| Rolling mean drifts upward | Changing level or trend | Whether the trend is deterministic or a unit root |
| Rolling standard deviation grows | Changing variance | That a log transform is always the correct fix |
| ACF remains close to 1 for many lags | Strong persistence, trend, or unit root | Non-stationarity by itself |
| ACF peaks at 24 and 168 | Daily and weekly seasonality in hourly data | That all other structure is stationary |
| Sudden permanent level shift | Structural break or new regime | That ordinary differencing alone will solve it |
Transformation question
What remains after subtracting yesterday from today? What remains after subtracting the same weekday last week?
Differencing: model movement instead of level
First differencing subtracts the previous observation from the current observation:
The original series answers “where are we?” The differenced series answers “how far did we move?” A level can keep rising while its movement follows a much simpler and more stable process.
Escalator intuition: a person standing on an upward-moving escalator gets higher above the ground every second. Their height is non-stationary, but the distance travelled each second may be nearly constant. Differencing removes the accumulated height and exposes the step-by-step movement.
Why forecasting models use differencing
Remove a changing level
A trend makes early and late observations live around different means. Differences can turn a rising level into stable average growth.
Remove accumulated shocks
In a random walk, every shock permanently changes the level. Differencing recovers the individual innovations.
Reveal genuine memory
A trending series can show high autocorrelation merely because adjacent observations share the same trend. Differencing helps separate movement from shared time position.
Stabilize model parameters
ARMA-style relationships are easier to estimate and interpret when the modeled series has a stable level and lag structure.
The random-walk derivation
In a random walk, today's level equals yesterday's level plus a new shock:
Shocks accumulate, so the level has no fixed long-run center. Subtracting yesterday from today gives:
The unstable accumulated level disappears. If the shocks \(\varepsilon_t\) form a stationary process, the first-differenced series is stationary. This is the central idea behind the \(I\), or integrated, part of ARIMA.
First difference
Compares neighboring observations and can remove a trend or unit root.
Seasonal difference
Compares matching seasonal positions \(m\) periods apart.
Here \(B\) is the backshift operator: \(BY_t=Y_{t-1}\).
The raw teaching series contains trend and weekly seasonality. First differencing reduces trend but retains weekly alternation. Seasonal differencing compares matching weekdays and removes most of the designed weekly structure; unusual events remain visible.
A small first-difference example
Suppose demand is \(100,108,116,124,132\). Its mean keeps moving upward, but its changes are stable:
The simpler statement is now visible: demand grows by approximately 8 units per period.
Forecast changes, then reconstruct the level
A model fitted to differences predicts \(\widehat{\Delta Y}_{t+h}\), not the final level. For one step ahead:
If the latest demand is 124 and the predicted change is 7, the level forecast is \(124+7=131\). For two steps, accumulate both predicted changes:
Differencing changes the prediction target: the last observed level acts as an anchor, and forecasted changes build the future path from that anchor.
Seasonal differencing: compare like with like
First differencing compares today with yesterday. Seasonal differencing compares today with the corresponding point in the previous cycle:
Suppose two weeks of daily demand are:
With \(m=7\), every seasonal difference is 5. The large weekday-weekend shape disappears, leaving the simpler week-over-week increase.
| Operation | Comparison | Typical purpose |
|---|---|---|
| First difference \((1-B)Y_t\) | Current value minus previous value | Remove trend or a nonseasonal unit root |
| Seasonal difference \((1-B^m)Y_t\) | Current value minus the matching seasonal value | Remove seasonal level or a seasonal unit root |
| Both \((1-B)(1-B^m)Y_t\) | Seasonally difference, then difference again | Handle both sources only when diagnostics justify it |
For hourly bike demand, \(m=24\) compares the same hour yesterday, while \(m=168\) compares the same hour and weekday last week.
How do we know differencing helped?
- The transformed plot has a more stable level and spread.
- The ACF no longer stays extremely high across many lags.
- The intended trend or seasonal peaks are reduced.
- ADF and KPSS evidence becomes more consistent with stationarity.
- Rolling-origin forecast accuracy improves, not merely the statistical-test result.
Do not difference automatically. Over-differencing can amplify noise, discard useful level information, widen long-horizon uncertainty, and create artificial negative lag-1 autocorrelation. A smooth deterministic trend may be better represented directly with a time feature or trend component. Difference only when modeling changes creates a more stable and useful forecasting mechanism.
Test question
If an ADF test rejects a unit root, has it proven that every stationarity condition is satisfied?
ADF test: does a shock persist forever?
Begin with an AR(1) process:
Subtract \(Y_{t-1}\) from both sides and define \(\gamma=\rho-1\):
Unit root
If \(\rho=1\), then \(\gamma=0\). Shocks accumulate permanently and the level behaves like a random walk.
Mean reversion
If \(\gamma<0\), a high previous level predicts a negative change and a low level predicts a positive change. The process is pulled back.
ADF intuition: test whether yesterday's level creates a restoring force in today's change. No restoring force is evidence for a unit root.
Why “augmented”?
The test adds lagged changes to absorb short-term autocorrelation:
| Term | Role |
|---|---|
| \(\alpha\) | Optional constant or drift |
| \(\beta t\) | Optional deterministic trend |
| \(\gamma Y_{t-1}\) | Term tested for a unit root |
| \(\Delta Y_{t-1},\ldots,\Delta Y_{t-p}\) | Lagged changes that remove residual autocorrelation |
ADF null
\(H_0:\gamma=0\)
The series contains a unit root.
ADF alternative
\(H_1:\gamma<0\)
The series is stationary under the selected constant/trend specification.
Suppose the ADF statistic is \(-4.21\), its 5% critical value is \(-2.86\), and \(p=0.0007\). Because the statistic is more negative than the critical value and \(p<0.05\), reject the unit-root null.
Careful wording: “The test provides evidence against a unit root under the chosen lag and deterministic-term specification.” It does not prove constant variance, remove seasonality, detect every structural break, or guarantee good forecasts. Failing to reject \(H_0\) also does not prove a unit root; the test may have low power.
Use ADF and KPSS as complementary evidence
ADF test
\(H_0\): the series contains a unit root.
A small \(p\)-value supports rejecting that unit-root explanation.
KPSS test
\(H_0\): the series is level-stationary or trend-stationary, depending on the test specification.
A small \(p\)-value is evidence against stationarity.
| ADF | KPSS | Practical reading |
|---|---|---|
| Reject unit root | Do not reject stationarity | Both are consistent with stationarity. |
| Do not reject unit root | Reject stationarity | Both are consistent with non-stationarity. |
| Mixed or inconclusive result | Inspect trend terms, lag choices, sample size, seasonality, and structural breaks. Do not choose a conclusion from the preferred test alone. | |
A test does not replace a plot. Results depend on lag selection, whether a constant or trend is included, sample size, and breaks. Stationarity is a working model assumption supported by converging evidence, not a permanent label attached to a dataset.
Does every forecasting model require stationarity?
| Model family | How it handles changing structure |
|---|---|
| ARMA | Usually applied to a stationary series. |
| ARIMA | The \(I\) component differences a non-stationary level before modeling AR and MA dependence. |
| Exponential smoothing | Represents evolving level, trend, and seasonality directly; the raw series need not be stationary. |
| Machine-learning models | Can use time, lag, rolling, and calendar features, but still fail when training relationships do not survive into deployment. |
The real objective: do not make a series stationary merely to satisfy a test. Make the mechanism being modeled stable enough that relationships learned from historical data remain useful at future forecast origins.