1. What Is Smoothing?

Time-series data often contains:

Smoothing tries to reduce the effect of short-term noise so that the underlying level, trend, or seasonal pattern becomes easier to identify.

Suppose daily demand is:

\[100,\ 104,\ 98,\ 103,\ 97,\ 105,\ 101\]

The values fluctuate, but the underlying level appears to be around 101.

Smoothing asks:

What is the current underlying level after ignoring some of the random movement?

2. Why Not Use the Latest Observation Directly?

If today’s demand is 130 because of an unusual event, predicting 130 for tomorrow may be too reactive.

On the other hand, using the average of the entire historical dataset may respond too slowly when demand genuinely changes.

We need a balance between:

This is the central idea behind exponential smoothing.

3. Moving-Average Intuition

A simple approach is to average the most recent \(w\) observations.

For a three-day moving average:

\[M_t=\frac{Y_t+Y_{t-1}+Y_{t-2}}{3}\]

If recent demand is:

\[100,\ 110,\ 105\]

then:

\[M_t=\frac{100+110+105}{3}=105\]

This reduces noise, but it has two limitations:

  1. Every observation inside the window receives equal weight.
  2. Observations suddenly receive zero weight when they leave the window.

Exponential smoothing uses weights that decrease gradually instead.

4. Simple Exponential Smoothing

Simple exponential smoothing maintains one quantity:

\[\ell_t=\text{estimated current level}\]

The update equation is:

\[\ell_t=\alpha Y_t+(1-\alpha)\ell_{t-1}\]

where:

5. Weighted-Average Intuition

The new level is a compromise between:

\[ \underbrace{\ell_t}_{\text{new belief}} = \underbrace{\alpha Y_t}_{\text{new evidence}} + \underbrace{(1-\alpha)\ell_{t-1}}_{\text{old belief}} \]

If \(\alpha=0.3\), the new level uses:

6. Walk-Toward-the-Observation Intuition

The same equation can be rewritten as:

\[ \ell_t = \ell_{t-1} + \alpha(Y_t-\ell_{t-1}) \]

The quantity:

\[Y_t-\ell_{t-1}\]

is the latest forecast error.

Therefore:

Start at the old level and move \(\alpha\) of the way toward the new observation.

7. Numerical Example

Suppose:

\[\ell_{t-1}=100,\qquad Y_t=120,\qquad \alpha=0.3\]

The error is:

\[120-100=20\]

Move 30% of that distance:

\[\ell_t=100+0.3(20)=106\]

The new observation was 120, but the estimated level moves only from 100 to 106.

The model treats part of the jump as possible noise.

If the next observation is 90:

\[\ell_{t+1}=0.3(90)+0.7(106)=101.2\]

The estimate moves downward but does not immediately jump to 90.

8. Meaning of \(\alpha\)

Small \(\alpha\)

For example:

\[\alpha=0.1\]

Large \(\alpha\)

For example:

\[\alpha=0.9\]

The choice is a bias-responsiveness tradeoff.

9. Why Is It Called “Exponential”?

Repeatedly substitute the previous level:

\[\ell_t=\alpha Y_t+(1-\alpha)\ell_{t-1}\]

Since:

\[\ell_{t-1}=\alpha Y_{t-1}+(1-\alpha)\ell_{t-2}\]

we obtain:

\[ \ell_t = \alpha Y_t + \alpha(1-\alpha)Y_{t-1} + \alpha(1-\alpha)^2Y_{t-2} +\cdots \]

The weights are:

\[\alpha,\quad\alpha(1-\alpha),\quad\alpha(1-\alpha)^2,\ldots\]

They decline geometrically, or exponentially, as observations become older.

Older observations are not discarded abruptly. Their influence gradually becomes smaller.

10. Forecasting With SES

Simple exponential smoothing has only a level state. It has no trend or seasonal state.

Therefore, all future forecasts are equal to the latest estimated level:

\[\widehat Y_{t+h\mid t}=\ell_t\]

for:

\[h=1,2,3,\ldots\]

If:

\[\ell_t=106\]

then:

\[ \widehat Y_{t+1\mid t} = \widehat Y_{t+2\mid t} = \widehat Y_{t+3\mid t} = 106 \]

SES produces a flat forecast path.

Use SES when the series has:

11. Why SES Cannot Handle Trend

Suppose demand is:

\[100,\ 110,\ 120,\ 130,\ 140\]

SES updates the level, but its future forecasts remain flat.

It may estimate the current level near 130 or 135, but it does not know that demand is increasing by approximately 10 every period.

To model that movement, we need a separate trend state.

12. Holt’s Method

Holt’s method maintains two states:

\[\ell_t=\text{current level}\]
\[b_t=\text{current trend}\]

The equations are:

\[\ell_t=\alpha Y_t+(1-\alpha)(\ell_{t-1}+b_{t-1})\]
\[b_t=\beta(\ell_t-\ell_{t-1})+(1-\beta)b_{t-1}\]

The forecast is:

\[\widehat Y_{t+h\mid t}=\ell_t+hb_t\]

13. Holt’s Method in Small Steps

Suppose:

\[\ell_{t-1}=100\]
\[b_{t-1}=5\]
\[Y_t=112\]
\[\alpha=0.4,\qquad\beta=0.3\]

Step 1: Predict today before seeing it

The old level was 100 and the old trend was 5:

\[100+5=105\]

The model expected today’s value to be 105.

Step 2: Update the level

Blend the actual observation with the projected level:

\[\ell_t=0.4(112)+0.6(105)=107.8\]

Step 3: Observe the new level change

\[\ell_t-\ell_{t-1}=107.8-100=7.8\]

The newly observed growth is 7.8.

Step 4: Update the trend

Blend the new growth with the old trend:

\[b_t=0.3(7.8)+0.7(5)=5.84\]

Step 5: Forecast

One step ahead:

\[\widehat Y_{t+1\mid t}=107.8+5.84=113.64\]

Two steps ahead:

\[\widehat Y_{t+2\mid t}=107.8+2(5.84)=119.48\]

Unlike SES, Holt’s method produces a trending future path.

14. Meaning of \(\alpha\) and \(\beta\)

A large \(\beta\) allows the slope to change rapidly. A small \(\beta\) assumes that the trend evolves gradually.

15. Why Holt’s Method Cannot Handle Seasonality

Suppose bike demand has:

Holt’s method can model the overall level and growth, but it has no memory of which day of the week is being forecast.

We therefore add a seasonal state.

16. Holt-Winters Method

Additive Holt-Winters maintains three components:

\[\ell_t=\text{level}\]
\[b_t=\text{trend}\]
\[s_t=\text{seasonal effect}\]

For season length \(m\):

\[\ell_t=\alpha(Y_t-s_{t-m})+(1-\alpha)(\ell_{t-1}+b_{t-1})\]
\[b_t=\beta(\ell_t-\ell_{t-1})+(1-\beta)b_{t-1}\]
\[s_t=\gamma(Y_t-\ell_{t-1}-b_{t-1})+(1-\gamma)s_{t-m}\]

The forecast is:

\[\widehat Y_{t+h\mid t}=\ell_t+hb_t+s_{\text{matching season}}\]

17. Holt-Winters Intuition

Before updating the level:

  1. Remove the known seasonal effect.
  2. Estimate the underlying level.
  3. Update the trend.
  4. Estimate whether this seasonal position was stronger or weaker than usual.
  5. Store that updated seasonal effect for the next cycle.

For daily observations with weekly seasonality:

\[m=7\]

Each weekday has its own seasonal memory.

For hourly demand:

18. Small Holt-Winters Example

Suppose:

\[\ell_{t-1}=100,\qquad b_{t-1}=2\]

The seasonal effect for this weekday is:

\[s_{t-7}=15\]

Today’s observed demand is:

\[Y_t=120\]

Use:

\[\alpha=0.4,\qquad\beta=0.2,\qquad\gamma=0.3\]

Step 1: Remove seasonality

\[Y_t-s_{t-7}=120-15=105\]

The underlying nonseasonal value is approximately 105.

Step 2: Update level

\[\ell_t=0.4(105)+0.6(100+2)=103.2\]

Step 3: Update trend

\[b_t=0.2(103.2-100)+0.8(2)=2.24\]

Step 4: Update seasonality

The new seasonal evidence is:

\[120-100-2=18\]

Blend it with the previous seasonal effect:

\[s_t=0.3(18)+0.7(15)=15.9\]

The weekday effect changes from 15 to 15.9.

Step 5: Forecast the same weekday next week

\[\widehat Y_{t+7\mid t}=103.2+7(2.24)+15.9=134.78\]

19. Additive Versus Multiplicative Seasonality

Additive seasonality

\[Y_t=\ell_t+b_t+s_t+\varepsilon_t\]

Use it when seasonal effects remain approximately constant in units.

Example:

Every Saturday adds approximately 50 rentals.

Multiplicative seasonality

\[Y_t=(\ell_t+b_t)s_t+\varepsilon_t\]

Use it when seasonal effects grow with the level.

Example:

Saturday demand is approximately 30% above normal.

If the seasonal fluctuations increase as the series grows, multiplicative seasonality may be more appropriate.

20. The Three Methods as a Story

MethodWhat it remembersForecast shape
SESCurrent levelFlat
HoltLevel and trendTrending
Holt-WintersLevel, trend and seasonalityTrending and repeating

The progression is:

\[\text{Level}\rightarrow\text{Level + trend}\rightarrow\text{Level + trend + seasonality}\]

Final Intuition

Exponential smoothing is a controlled memory system.

Each new observation asks:

The smoothing parameters determine how much the model changes its beliefs after receiving new evidence.

Previous: Smoothing chapterNext: ARIMA Pre