Tomorrow is unavailable during training.
Time-series forecasting uses observations recorded in time order to estimate values that have not happened yet. Unlike ordinary regression, the future cannot be shuffled into training and every predictor must genuinely exist when the forecast is issued.
Opening problem
At 8 p.m., a bike-sharing operator must decide how many bikes to make available tomorrow. What could be predicted, what history may help, and what mistake would be expensive?
What is a forecasting problem?
We observe a sequence \(y_1,y_2,\ldots,y_t\) up to the current forecast origin \(t\). We use that history—and any other information genuinely available by \(t\)—to estimate a future value \(y_{t+h}\).
Forecast origin \(t\)
The moment the prediction is made. It creates the boundary between known and unknown.
Horizon \(h\)
How far ahead we predict: one hour, one day, or an entire week.
Frequency
Spacing between observations: hourly, daily, weekly, and so on.
Information set
All target history and external predictors available at the forecast origin.
Forecasting is not ordinary curve completion. Interpolation fills a gap surrounded by known points. Forecasting extends beyond the last known target, where uncertainty normally grows with the horizon.
The real dataset behind the story
We will connect the concepts to Kaggle's Bike Sharing Demand dataset. It contains hourly bike-rental observations spanning two years. The practical task is to forecast the total number of rentals for each requested hour.
Training file
The first 19 days of each month, including predictors and observed rental counts.
Competition test file
Days 20 through month-end, containing predictors but not rental outcomes.
Target
count: total hourly rentals.
Prediction unit
One row represents one hour. The datetime supplies hour, weekday, month and chronological order.
| Field | Meaning | Teaching role |
|---|---|---|
datetime | Hourly timestamp | Order, hour, weekday, month and lag alignment |
season, holiday, workingday | Calendar conditions | Known calendar predictors |
weather | Four ordered weather-condition categories | External demand driver |
temp, atemp | Temperature and feels-like temperature in Celsius | Continuous weather predictors |
humidity, windspeed | Relative humidity and wind speed | Continuous weather predictors |
casual, registered | Components of total rentals in the training file | Do not use as features when predicting count; they reveal the target |
count | Total hourly rentals | Forecast target |
Availability nuance: calendar fields are known in advance. Realized weather values may not be; a deployed forecast would need weather forecasts or scenarios. The competition supplies weather fields for test rows, but a production design must still ask how those values will be obtained.
How to decode the categorical fields
season
1 = spring, 2 = summer, 3 = fall, 4 = winter. Treat these as category labels rather than assuming the numeric gaps have physical meaning.
weather
1 = clear/few clouds; 2 = mist/cloudy; 3 = light snow or light rain/thunderstorm; 4 = severe rain, snow or fog.
holiday
Indicates whether the date is considered a holiday.
workingday
1 when the day is neither a weekend nor a holiday; otherwise 0.
The download contains train.csv, test.csv and sampleSubmission.csv. The submission requires two columns: datetime and the predicted count.
First visual prediction
What repeating and non-repeating patterns would you look for before fitting a model?
Teaching simulation, not raw Kaggle rows. We aggregate the story to a daily scale and construct gradual growth, a seven-day pattern, small noise, and two unusual days so every later component can be checked exactly. The coding notebook can then repeat the workflow on the hourly Kaggle data.
The central question: given observations available through time \(t\), what can we responsibly say about \(t+h\)?
The story we will build
Define the contract
Target, forecast origin, horizon, frequency, and information available at prediction time.
See the structure
Trend, seasonality, cycles, noise, outliers, and decomposition.
Earn complexity
Naive, seasonal naive, drift, and moving-average forecasts.
Measure memory
Lag features, autocorrelation, ACF, PACF, and availability.
Stabilize behavior
Differencing, seasonal differencing, transformations, and ADF.
Update state
Simple smoothing, Holt trend, and Holt-Winters seasonality.
Smoothing in 20 steps
A slower, spelled-out path from noisy observations to Holt-Winters and model selection.
ARIMA Pre
A step-by-step introduction to AR, integration, MA, orders, diagnostics, and seasonal ARIMA.
Model dependence
AR, MA, integration, seasonal ARIMA, and residual checks.
SARIMA Pre
Seasonal differencing, seasonal AR and MA terms, model notation, selection, and diagnostics.
Simulate deployment
Walk-forward evaluation, metrics, horizons, and intervals.
Engineer safely
Leakage-safe features, ML models, multi-step strategies, and monitoring.
What stays outside this first lecture
VAR, state-space derivations, Prophet, deep sequence models, hierarchical forecasting, intermittent-demand methods, and probabilistic forecasting deserve later sessions. This foundation first establishes correct reasoning and evaluation.