Separate repetition from movement and surprise.
A line plot is the first model of the problem. It reveals whether demand moves slowly, repeats on a calendar, changes variability, or experiences unusual events.
Pattern question
Which part should repeat every seven days? Which part should move slowly? Where would a festival, strike or station closure appear?
Four kinds of movement
Trend \(T_t\)
A sustained long-term increase or decrease in the local level. Population growth, system expansion or gradual adoption could raise bike demand over months.
Not every upward run is a trend: it should persist beyond a few adjacent points.
Seasonality \(S_t\)
A pattern tied to a known and fixed period \(m\): hour of day, day of week, or month of year. Commuter peaks can repeat every 24 hours; weekday structure every 168 hourly observations.
Its timing is predictable even when its size changes.
Cycle
A rise and fall without a fixed calendar period. Economic activity, fuel prices or long weather regimes may create multi-month cycles whose duration varies.
A cycle cannot be located by simply saying “every \(m\) observations.”
Remainder \(R_t\)
What remains after the estimated trend and seasonal pattern are removed. It contains random noise, measurement errors, unmodeled effects, outliers and structural breaks.
A large remainder is a question to investigate, not automatically “randomness.”
Important distinction: seasonality has a stable calendar period; a cycle has no fixed period; the remainder is whatever the chosen decomposition has not explained.
Read before labels
In the figure below, identify the slowly rising line, the seven-day pattern, and the two observations that the planned structure cannot explain.
One series, decomposed consistently
All four panels use the same 49 simulated days and aligned horizontal axis. The observed series was constructed as trend + weekly seasonal effect + remainder, so the equality can be checked point by point. The isolated spike and dip remain in the remainder.
Real decomposition is estimated. We do not normally know the true components. Moving averages, STL or model-based methods estimate them, and boundary values are often less reliable because fewer neighboring observations exist.
What would this mean in the hourly Kaggle data?
| Observed pattern | Likely interpretation | What to check |
|---|---|---|
| Morning and evening peaks on working days | Within-day seasonality and commuter behavior | Hour × working-day interaction |
| Different weekend shape | Weekly/calendar seasonality | Day of week, holiday and user type |
| Growth from 2011 to 2012 | Long-run trend | Year and expanding system usage |
| Demand collapse during heavy rain | Weather-driven external effect | Whether weather is available at forecast time |
| One unexplained extreme hour | Remainder or outlier | Event, outage, data error or new regime |
Multiple seasonalities: hourly demand may contain both a daily period \(m=24\) and a weekly period \(m=168\). A single seasonal component may be insufficient.
Scale question
If weekend demand grows from 20 above average to 60 above average as the system becomes busier, is additive seasonality still natural?
Additive or multiplicative?
Additive
Seasonal swings stay roughly constant in the target's units: every weekend adds about 40 rentals.
Multiplicative
Seasonality scales with the level: weekend demand is about 30% above the current level.
A logarithm turns multiplicative relationships into additive ones, but it requires positive values and changes how forecasts are interpreted after transforming back.
Before modeling: verify frequency, duplicates, missing timestamps, time zone, daylight-saving changes, outliers, and whether zeros are real. A clean-looking plot can still hide a broken time index.