ARIMA Pre

1. Start With the Problem

Suppose monthly demand is:

\[100,\ 108,\ 115,\ 121,\ 130,\ 138,\ldots\]

The series is increasing, so predicting with a fixed historical mean will not work well.

After removing the growth through differencing:

\[\Delta Y_t=Y_t-Y_{t-1}\]

we get:

\[8,\ 7,\ 6,\ 9,\ 8,\ldots\]

These changes look more stable. However, they may still contain predictable relationships:

ARIMA models these relationships.

ARIMA models a stationary version of the series using its past values and past forecast errors.

2. What Does ARIMA Mean?

ARIMA stands for:

\[\boxed{\text{ARIMA}(p,d,q)}\]
ComponentMeaningMain question
ARAutoRegressiveDo previous values help predict the current value?
IIntegratedHow many times should we difference the series?
MAMoving AverageDo previous forecast errors help predict the current value?

The parameters are:

Part 1: AR

3. Autoregression Intuition

“Auto” means self. “Regression” means predicting a value using other values.

An autoregressive model predicts a series using its own past:

\[Y_t=c+\phi_1Y_{t-1}+\phi_2Y_{t-2}+\cdots+\phi_pY_{t-p}+\varepsilon_t\]

This is called an AR(\(p\)) model.

For an AR(1) model:

\[Y_t=c+\phi Y_{t-1}+\varepsilon_t\]

Example

Suppose:

\[Y_t=20+0.7Y_{t-1}+\varepsilon_t\]

If yesterday’s demand was 100 and we ignore the unknown future error:

\[\widehat Y_t=20+0.7(100)=90\]

The model says:

Today’s demand is related to yesterday’s demand, but it does not simply copy it.

4. Meaning of the AR Coefficient

Consider:

\[Y_t=c+\phi Y_{t-1}+\varepsilon_t\]
\[|\phi|<1\]

Long-run mean

For a stationary AR(1), the long-run mean is:

\[\mu=\frac{c}{1-\phi}\]

For:

\[Y_t=20+0.7Y_{t-1}+\varepsilon_t\]

the long-run mean is:

\[\mu=\frac{20}{1-0.7}=66.67\]

If the series is above 66.67, it tends to move downward. If it is below 66.67, it tends to move upward.

This is mean reversion.

Part 2: I

5. Why Do We Need Integration?

AR and MA models assume that the modeled series is stationary.

But suppose the series follows a random walk:

\[Y_t=Y_{t-1}+\varepsilon_t\]

Every new error permanently changes the level. The series has no stable long-run mean.

Taking the first difference gives:

\[\Delta Y_t=Y_t-Y_{t-1}=\varepsilon_t\]

The unstable accumulated level disappears, leaving a potentially stationary series.

The “Integrated” component means:

  1. Difference the series before modeling.
  2. Forecast the differences.
  3. Integrate, or accumulate, those forecasts to recover the original level.

6. Meaning of \(d\)

For ARIMA(\(p,d,q\)):

First difference:

\[\Delta Y_t=Y_t-Y_{t-1}\]

Second difference:

\[\Delta^2Y_t=\Delta Y_t-\Delta Y_{t-1}\]

Most practical series use \(d=0\) or \(d=1\). Using unnecessary differences can amplify noise and create artificial negative autocorrelation.

7. Reconstructing the Forecast

Suppose the ARIMA model predicts:

\[\widehat{\Delta Y}_{t+1}=8\]

and the latest observed level is:

\[Y_t=130\]

Then:

\[\widehat Y_{t+1\mid t}=Y_t+\widehat{\Delta Y}_{t+1}=130+8=138\]

For two periods:

\[ \widehat Y_{t+2\mid t} = Y_t+ \widehat{\Delta Y}_{t+1\mid t}+ \widehat{\Delta Y}_{t+2\mid t} \]

The predicted changes are accumulated to reconstruct the future level.

Part 3: MA

8. Moving-Average Error Intuition

The MA component does not mean taking the average of recent observations.

An MA model uses previous forecast errors:

\[Y_t=c+\varepsilon_t+\theta_1\varepsilon_{t-1}+\cdots+\theta_q\varepsilon_{t-q}\]

For MA(1):

\[Y_t=c+\varepsilon_t+\theta\varepsilon_{t-1}\]

At forecasting time, the current error \(\varepsilon_t\) is unknown, but previous errors are known.

Example

Suppose:

\[Y_t=100+0.6\varepsilon_{t-1}+\varepsilon_t\]

Yesterday’s forecast was 90, but the actual value was 100:

\[\varepsilon_{t-1}=100-90=10\]

Therefore:

\[\widehat Y_t=100+0.6(10)=106\]

The model reasons:

I underpredicted yesterday. If that surprise has a short-lived effect, I should adjust today’s forecast upward.

9. AR Versus MA

AR asks:

Did the previous value contain useful information?
\[Y_t\leftarrow Y_{t-1},Y_{t-2},\ldots\]

MA asks:

Did the previous forecast surprise contain useful information?
\[Y_t\leftarrow\varepsilon_{t-1},\varepsilon_{t-2},\ldots\]
ARMA
Uses previous observationsUses previous forecast errors
Models persistence in valuesModels persistence in shocks
Past values are directly observedErrors are inferred after fitting

Putting Everything Together

10. The Complete ARIMA Model

First define the differenced series:

\[Z_t=(1-B)^dY_t\]

Here, \(B\) is the backshift operator:

\[BY_t=Y_{t-1}\]

For \(d=1\):

\[(1-B)Y_t=Y_t-Y_{t-1}\]

ARIMA then applies an ARMA model to \(Z_t\):

\[ Z_t = c+ \sum_{i=1}^{p}\phi_iZ_{t-i} + \varepsilon_t + \sum_{j=1}^{q}\theta_j\varepsilon_{t-j} \]

In plain language:

\[ \boxed{ \text{Current change} = \text{past changes} + \text{past surprises} + \text{new surprise} } \]

11. ARIMA(1,1,1) Example

ARIMA(1,1,1) means:

The model is:

\[\Delta Y_t=c+\phi\Delta Y_{t-1}+\theta\varepsilon_{t-1}+\varepsilon_t\]

Suppose:

\[c=2,\qquad \phi=0.6,\qquad \theta=0.4\]

The previous change was:

\[\Delta Y_t=10\]

and the previous error was:

\[\varepsilon_t=-3\]

The next predicted change is:

\[\widehat{\Delta Y}_{t+1}=2+0.6(10)+0.4(-3)=6.8\]

If the latest observed level is 150:

\[\widehat Y_{t+1\mid t}=150+6.8=156.8\]

Notice the complete process:

  1. Predict the next change.
  2. Add the predicted change to the latest level.

12. Important Special Cases

ModelInterpretation
ARIMA(0,0,0)White noise around a constant
ARIMA(1,0,0)AR(1)
ARIMA(0,0,1)MA(1)
ARIMA(\(p\),0,\(q\))ARMA(\(p,q\))
ARIMA(0,1,0)Random walk
ARIMA(0,1,0) with driftRandom walk with average growth
ARIMA(1,1,0)Current change depends on the previous change
ARIMA(0,1,1)Current change depends on the previous shock

A useful class question:

If ARIMA(0,1,0) is a random walk, what is its next forecast?

Because:

\[Y_t=Y_{t-1}+\varepsilon_t\]

and the expected future error is zero:

\[\widehat Y_{t+1\mid t}=Y_t\]

It produces the naive forecast.

Choosing \(p,d,q\)

13. Choose \(d\) First

Start by asking whether the series needs differencing.

Use:

A slowly decaying ACF can indicate non-stationarity, but it is not sufficient by itself.

Choose the smallest \(d\) that creates a reasonably stable series.

14. ACF and PACF Hints

For an ideal stationary process:

PatternPossible model
PACF cuts off after lag \(p\), ACF decaysAR(\(p\))
ACF cuts off after lag \(q\), PACF decaysMA(\(q\))
Both gradually decayARMA model may be needed

These are diagnostic hints, not strict rules for real data.

A good workflow is:

  1. Use ACF/PACF to propose a few small models.
  2. Fit those models.
  3. Compare AIC/BIC.
  4. Compare rolling-origin forecast performance.
  5. Check residuals.

Avoid searching over very large \(p\) and \(q\) without justification.

Model Diagnostics

15. What Should Good Residuals Look Like?

After fitting ARIMA:

\[e_t=Y_t-\widehat Y_{t\mid t-1}\]

The residuals should have:

Use:

If residual autocorrelation remains, the model has left predictable information behind.

Seasonal ARIMA

16. What About Seasonality?

For seasonal data, we use:

\[\operatorname{SARIMA}(p,d,q)(P,D,Q)_m\]

The seasonal parameters are:

For daily observations with weekly seasonality:

\[m=7\]

Seasonal differencing is:

\[\Delta_7Y_t=Y_t-Y_{t-7}\]

A seasonal AR term relates the current value to a matching previous season. A seasonal MA term uses a forecast error from a previous season.

17. ARIMA Versus Exponential Smoothing

Exponential smoothingARIMA
Models evolving level, trend, and seasonalityModels autocorrelation in a stationary or differenced series
State-based intuitionLag-and-error-based intuition
Strong for changing local structureStrong for systematic temporal dependence
Holt-Winters represents seasonality through seasonal statesSARIMA represents seasonality through seasonal lags

Neither is universally superior. Compare them using the same rolling-origin validation setup.

18. When Is ARIMA Useful?

ARIMA is especially useful when:

ARIMA may struggle when:

Final Intuition

ARIMA can be remembered as three operations:

\[\boxed{\text{Stabilize}\rightarrow\text{remember values}\rightarrow\text{correct using surprises}}\]

The central idea is:

Once the changing level has been handled, can past movements and past surprises explain what happens next?
Previous: SmoothingNext: ARIMA