Machine learning sees columns, not time.
Linear models, forests and boosting can forecast after we encode temporal context. The feature table—and when each value becomes available—is part of the model.
Feature-engineering question
Can Random Forest discover “same weekday last week” if the table contains only today's row number and target?
A forecast-ready row
| Origin | Lag 1 | Lag 7 | Rolling mean 7 | Weekday | Weather forecast | Target |
|---|---|---|---|---|---|---|
| Sunday 8 p.m. | Sunday demand | Previous Monday | Previous seven complete days | Monday | Forecast issued Sunday | Monday demand |
Known future features
Calendar, planned price, scheduled event, committed promotion.
Unknown future features
Observed weather, future demand, unplanned outage. Use forecasts or scenarios, not realized values.
Common leakage: computing rolling statistics on the full dataset before splitting; filling past missing values from future observations; scaling on all dates; selecting features using test performance.
Fourteen-day question
After predicting tomorrow, how will a lag-based model obtain the lag-1 value needed to predict the day after tomorrow?
Multi-step strategies
Recursive
Fit one one-step model and feed predictions back as future lags. Simple, but errors accumulate.
Direct
Fit a separate model for each horizon. Avoids feedback, but uses more models and data.
Multi-output
Predict the whole path together. Can learn horizon relationships but is more complex.
Match evaluation to the strategy: recursive models must be evaluated recursively. Replacing predicted lags with actual future values during testing gives an unrealistically easy task.
What different model families contribute
| Model | Strength | Watch for |
|---|---|---|
| Linear or regularized regression | Transparent lag and calendar effects | Manual nonlinearities; scaling for regularization |
| Decision Tree | Rules and interactions | High variance; stepwise predictions |
| Random Forest | Stable nonlinear baseline | Poor extrapolation beyond observed targets |
| Gradient Boosting | Strong tabular interactions and residual correction | Tuning, overfit, and no automatic time awareness |
| ETS / ARIMA | Purpose-built temporal structure and statistical intervals | Specification and changing regimes |
Tree ensembles interpolate patterns in their training target range; they do not naturally continue a deterministic upward trend. Include trend-related features, transform the task, or choose a model whose structure supports extrapolation.
Diagnosis question
A model has low average MAE but consistently underpredicts weekends. Is the work finished?
The defensible workflow
Residual questions
- Is mean error near zero, overall and by segment?
- Does residual ACF contain predictable structure?
- Does variance increase with fitted demand?
- Do errors deteriorate with forecast horizon?
- Are intervals calibrated and operationally useful?
Deployment questions
- When do actual labels arrive?
- How often will the model refit?
- What happens when required predictors are unavailable?
- Will drift monitoring compare like-for-like seasons?
- What baseline triggers investigation or rollback?
Final principle: the best forecasting model is the simplest complete workflow that beats a relevant baseline under a deployment-matched backtest and leaves no important predictable error.
Final retrieval
Question 1
Why can random splitting make forecasting performance look too good?
Question 2
What does seasonal naive test?
Question 3
How do AR and MA differ?
Question 4
What must a residual ACF ideally show?