Part 5

Model a stable mechanism, not a moving target.

Stationarity asks whether the statistical behavior observed in one period remains a useful description of another period. It is the bridge between “this happened in the past” and “this relationship may repeat in the future.”

Window-shift question

Imagine hiding the date labels and showing three different months from the series. Would their level, spread, and short-term movement look as though they came from the same process?

What is stationarity?

A time series is stationary when its probability-generating mechanism does not change as time moves forward. The values may rise and fall, but the rules governing those movements remain stable.

Classroom intuition: slide a window along the timeline. A stationary series should show broadly similar center, variability, and dependence in each window. It does not need to repeat the exact same values.

Stationary-looking

\(98,103,100,96,102,101,99,\ldots\)

Values fluctuate around approximately 100 with a similar spread throughout.

Non-stationary trend

\(100,108,116,124,132,140,\ldots\)

The typical level increases, so an early-window mean no longer describes a later window.

Three examples used to diagnose stationarity The first series fluctuates around a stable mean, the second has an upward trend, and the third has variance that increases over time. Stable leveland spreadChangingmeanChangingvariance Earlier timeLater time

The first series changes value without changing its behavior. The second changes its mean; the third keeps a similar center but its fluctuations grow. Only the first is stationary-looking.

Meaning question

Does stationary mean “constant,” “without noise,” or “perfectly predictable”?

Weak stationarity, one condition at a time

\[\mathbb E[Y_t]=\mu,\qquad \operatorname{Var}(Y_t)=\sigma^2,\qquad \operatorname{Cov}(Y_t,Y_{t-k})=\gamma_k\]

1. Stable mean

\[\mathbb E[Y_t]=\mu\]

The expected center does not depend on calendar time. Demand cannot average 100 early in the series and 300 later while satisfying this condition.

2. Stable variance

\[\operatorname{Var}(Y_t)=\sigma^2\]

The typical size of fluctuations remains stable. A quiet first year followed by a highly volatile second year violates this condition.

3. Stable memory

\[\operatorname{Cov}(Y_t,Y_{t-k})=\gamma_k\]

Dependence is determined by the distance \(k\), not the date. The one-day relationship should be comparable in January and June.

What this explains: a relationship estimated from historical windows can be reused at a future forecast origin. If the mean, variance, or lag relationship keeps changing, one fixed model is trying to learn several different processes.

Weak versus strict stationarity: strict stationarity requires the complete joint distribution to remain unchanged after shifting time. In practical forecasting, weak stationarity is more commonly discussed because it focuses on mean, variance, and covariance.

Generalization question

A model learns that demand usually returns toward 100. If the business expands and the new normal becomes 200, should that old relationship still be trusted?

Why stationarity matters in forecasting

A forecasting model learns relationships from earlier observations and applies them at a later forecast origin. This transfer from past to future is credible only when the learned mechanism remains reasonably stable.

\[\text{historical relationship}\quad\Longrightarrow\quad\text{future relationship}\]

Core intuition: stationarity does not make the future certain. It makes the past statistically relevant. Without some form of stability, excellent historical fit may describe a regime that no longer exists.

1. Models can generalize

A stable mean, variance, and lag structure allow relationships estimated in training windows to remain useful in validation and deployment.

2. Parameters keep their meaning

An AR coefficient should describe the same type of dependence throughout the series. If dependence changes with time, one fixed coefficient becomes an average of incompatible regimes.

3. ACF and PACF are interpretable

Classical autocorrelation assumes that dependence at lag k is comparable wherever it is measured. Trend and changing seasonality can create correlations that reflect shared time movement rather than useful memory.

4. Uncertainty is more credible

If error variability is stable, historical residuals can inform future prediction intervals. When variance grows, intervals learned from quiet periods become too narrow.

5. Residual checks become meaningful

After a model removes predictable structure, its residuals should behave like a stable, patternless process. Remaining trend, seasonality, or changing spread signals unfinished modeling.

6. Forecast errors reveal change

When a previously stable model develops persistent bias, it can indicate drift, a structural break, or a new operating regime rather than ordinary random error.

A model trained around the wrong level

Suppose historical demand fluctuates around 100 and we estimate:

\[Y_t-100=0.7(Y_{t-1}-100)+\varepsilon_t\]

If yesterday's demand is 120, the one-step forecast is:

\[\widehat Y_{t+1}=100+0.7(120-100)=114\]

The model expects deviations to move back toward 100. But if expansion permanently changes the normal level to 200, that mean-reversion target is obsolete. The same fitted equation will systematically underforecast.

Important qualification: the raw series does not always need to be stationary. Exponential smoothing can model an evolving level, regression can model trend and calendar effects, and ARIMA can difference the series. What must remain stable is the mechanism left for the model to learn after those components are represented.

Examples and non-examples

SeriesStationary?Reason
Noise fluctuating around 100 with stable spreadPossiblyIts level and variance can remain stable even though individual values are unpredictable.
Bike demand growing as the service expandsNoThe mean changes through time because of trend.
Hourly demand with a repeated morning peakNot in raw formThe expected value depends on the hour. Seasonal adjustment or seasonal differencing may reveal a stable remainder.
Sales whose fluctuations grow with sales levelNoThe variance changes; a logarithm may stabilize proportional variation.
Random walk: \(Y_t=Y_{t-1}+\varepsilon_t\)NoEvery shock permanently changes the level, and \(\operatorname{Var}(Y_t)\) grows with time.
First difference of a random walk: \(\Delta Y_t=\varepsilon_t\)Often yesDifferencing removes the accumulated level and leaves the innovations.

Seasonality nuance: deterministic seasonality is predictable rather than random, but it still makes the unconditional mean depend on time position. We can model the seasonal state directly or remove it before fitting a model that assumes stationarity.

Diagnosis question

If the first and second half have similar averages, is that enough to declare the series stationary?

How to identify stationarity

  1. Plot the values in time order. Look for trend, repeating seasonal level, growing spread, sudden breaks, and long periods at different levels. A time plot is more informative than a histogram because it preserves order.
  2. Compare time windows. Split the series into early, middle, and late windows. Compare their mean and standard deviation. Large systematic differences suggest changing level or variance.
  3. Plot rolling statistics. A rolling mean exposes local level; a rolling standard deviation exposes local spread. Use them as visual diagnostics, not formal proof, because their appearance depends on window length.
  4. Inspect the ACF. A trending or unit-root series often has large positive autocorrelations that decay very slowly. Seasonal peaks indicate that the mean depends on calendar position. A stationary series usually loses correlation more quickly, although persistent stationary processes can still decay slowly.
  5. Use statistical tests. ADF tests a unit-root null; KPSS tests a stationarity null. Their hypotheses differ, so using both can expose uncertainty rather than forcing one test to make the entire decision.
  6. Repeat after transformation. If you apply a log, first difference, or seasonal difference, replot the result and rerun the diagnostics. The transformed series, not the original one, is what the stationary model will see.
Diagnostic observationWhat it suggestsWhat it does not prove
Rolling mean drifts upwardChanging level or trendWhether the trend is deterministic or a unit root
Rolling standard deviation growsChanging varianceThat a log transform is always the correct fix
ACF remains close to 1 for many lagsStrong persistence, trend, or unit rootNon-stationarity by itself
ACF peaks at 24 and 168Daily and weekly seasonality in hourly dataThat all other structure is stationary
Sudden permanent level shiftStructural break or new regimeThat ordinary differencing alone will solve it

Transformation question

What remains after subtracting yesterday from today? What remains after subtracting the same weekday last week?

Differencing: model movement instead of level

First differencing subtracts the previous observation from the current observation:

\[\Delta Y_t=Y_t-Y_{t-1}\]

The original series answers “where are we?” The differenced series answers “how far did we move?” A level can keep rising while its movement follows a much simpler and more stable process.

Escalator intuition: a person standing on an upward-moving escalator gets higher above the ground every second. Their height is non-stationary, but the distance travelled each second may be nearly constant. Differencing removes the accumulated height and exposes the step-by-step movement.

Why forecasting models use differencing

Remove a changing level

A trend makes early and late observations live around different means. Differences can turn a rising level into stable average growth.

Remove accumulated shocks

In a random walk, every shock permanently changes the level. Differencing recovers the individual innovations.

Reveal genuine memory

A trending series can show high autocorrelation merely because adjacent observations share the same trend. Differencing helps separate movement from shared time position.

Stabilize model parameters

ARMA-style relationships are easier to estimate and interpret when the modeled series has a stable level and lag structure.

The random-walk derivation

In a random walk, today's level equals yesterday's level plus a new shock:

\[Y_t=Y_{t-1}+\varepsilon_t\]

Shocks accumulate, so the level has no fixed long-run center. Subtracting yesterday from today gives:

\[\Delta Y_t=Y_t-Y_{t-1}=\varepsilon_t\]

The unstable accumulated level disappears. If the shocks \(\varepsilon_t\) form a stationary process, the first-differenced series is stationary. This is the central idea behind the \(I\), or integrated, part of ARIMA.

First difference

\[\Delta Y_t=Y_t-Y_{t-1}=(1-B)Y_t\]

Compares neighboring observations and can remove a trend or unit root.

Seasonal difference

\[\Delta_mY_t=Y_t-Y_{t-m}=(1-B^m)Y_t\]

Compares matching seasonal positions \(m\) periods apart.

Here \(B\) is the backshift operator: \(BY_t=Y_{t-1}\).

First differenceSeasonal difference

The raw teaching series contains trend and weekly seasonality. First differencing reduces trend but retains weekly alternation. Seasonal differencing compares matching weekdays and removes most of the designed weekly structure; unusual events remain visible.

A small first-difference example

Suppose demand is \(100,108,116,124,132\). Its mean keeps moving upward, but its changes are stable:

\[\Delta Y=[108-100,\ 116-108,\ 124-116,\ 132-124]=[8,8,8,8]\]

The simpler statement is now visible: demand grows by approximately 8 units per period.

Forecast changes, then reconstruct the level

A model fitted to differences predicts \(\widehat{\Delta Y}_{t+h}\), not the final level. For one step ahead:

\[\widehat Y_{t+1\mid t}=Y_t+\widehat{\Delta Y}_{t+1\mid t}\]

If the latest demand is 124 and the predicted change is 7, the level forecast is \(124+7=131\). For two steps, accumulate both predicted changes:

\[\widehat Y_{t+2\mid t}=Y_t+\widehat{\Delta Y}_{t+1\mid t}+\widehat{\Delta Y}_{t+2\mid t}\]

Differencing changes the prediction target: the last observed level acts as an anchor, and forecasted changes build the future path from that anchor.

Seasonal differencing: compare like with like

First differencing compares today with yesterday. Seasonal differencing compares today with the corresponding point in the previous cycle:

\[\Delta_mY_t=Y_t-Y_{t-m}\]

Suppose two weeks of daily demand are:

\[\begin{aligned}\text{Week 1: }&[100,110,120,130,140,170,160]\\\text{Week 2: }&[105,115,125,135,145,175,165]\end{aligned}\]

With \(m=7\), every seasonal difference is 5. The large weekday-weekend shape disappears, leaving the simpler week-over-week increase.

OperationComparisonTypical purpose
First difference \((1-B)Y_t\)Current value minus previous valueRemove trend or a nonseasonal unit root
Seasonal difference \((1-B^m)Y_t\)Current value minus the matching seasonal valueRemove seasonal level or a seasonal unit root
Both \((1-B)(1-B^m)Y_t\)Seasonally difference, then difference againHandle both sources only when diagnostics justify it

For hourly bike demand, \(m=24\) compares the same hour yesterday, while \(m=168\) compares the same hour and weekday last week.

How do we know differencing helped?

Do not difference automatically. Over-differencing can amplify noise, discard useful level information, widen long-horizon uncertainty, and create artificial negative lag-1 autocorrelation. A smooth deterministic trend may be better represented directly with a time feature or trend component. Difference only when modeling changes creates a more stable and useful forecasting mechanism.

Test question

If an ADF test rejects a unit root, has it proven that every stationarity condition is satisfied?

ADF test: does a shock persist forever?

Begin with an AR(1) process:

\[Y_t=\rho Y_{t-1}+\varepsilon_t\]

Subtract \(Y_{t-1}\) from both sides and define \(\gamma=\rho-1\):

\[\Delta Y_t=(\rho-1)Y_{t-1}+\varepsilon_t=\gamma Y_{t-1}+\varepsilon_t\]

Unit root

If \(\rho=1\), then \(\gamma=0\). Shocks accumulate permanently and the level behaves like a random walk.

Mean reversion

If \(\gamma<0\), a high previous level predicts a negative change and a low level predicts a positive change. The process is pulled back.

ADF intuition: test whether yesterday's level creates a restoring force in today's change. No restoring force is evidence for a unit root.

Why “augmented”?

The test adds lagged changes to absorb short-term autocorrelation:

\[\Delta Y_t=\alpha+\beta t+\gamma Y_{t-1}+\sum_{i=1}^{p}\delta_i\Delta Y_{t-i}+\varepsilon_t\]
TermRole
\(\alpha\)Optional constant or drift
\(\beta t\)Optional deterministic trend
\(\gamma Y_{t-1}\)Term tested for a unit root
\(\Delta Y_{t-1},\ldots,\Delta Y_{t-p}\)Lagged changes that remove residual autocorrelation

ADF null

\(H_0:\gamma=0\)

The series contains a unit root.

ADF alternative

\(H_1:\gamma<0\)

The series is stationary under the selected constant/trend specification.

Suppose the ADF statistic is \(-4.21\), its 5% critical value is \(-2.86\), and \(p=0.0007\). Because the statistic is more negative than the critical value and \(p<0.05\), reject the unit-root null.

Careful wording: “The test provides evidence against a unit root under the chosen lag and deterministic-term specification.” It does not prove constant variance, remove seasonality, detect every structural break, or guarantee good forecasts. Failing to reject \(H_0\) also does not prove a unit root; the test may have low power.

Use ADF and KPSS as complementary evidence

ADF test

\(H_0\): the series contains a unit root.

A small \(p\)-value supports rejecting that unit-root explanation.

KPSS test

\(H_0\): the series is level-stationary or trend-stationary, depending on the test specification.

A small \(p\)-value is evidence against stationarity.

ADFKPSSPractical reading
Reject unit rootDo not reject stationarityBoth are consistent with stationarity.
Do not reject unit rootReject stationarityBoth are consistent with non-stationarity.
Mixed or inconclusive resultInspect trend terms, lag choices, sample size, seasonality, and structural breaks. Do not choose a conclusion from the preferred test alone.

A test does not replace a plot. Results depend on lag selection, whether a constant or trend is included, sample size, and breaks. Stationarity is a working model assumption supported by converging evidence, not a permanent label attached to a dataset.

Does every forecasting model require stationarity?

Model familyHow it handles changing structure
ARMAUsually applied to a stationary series.
ARIMAThe \(I\) component differences a non-stationary level before modeling AR and MA dependence.
Exponential smoothingRepresents evolving level, trend, and seasonality directly; the raw series need not be stationary.
Machine-learning modelsCan use time, lag, rolling, and calendar features, but still fail when training relationships do not survive into deployment.

The real objective: do not make a series stationary merely to satisfy a test. Make the mechanism being modeled stable enough that relationships learned from historical data remain useful at future forecast origins.

Previous: LagsNext: Smoothing