Part 8

Evaluate by repeatedly pretending to live in the past.

A trustworthy comparison reproduces the forecast origin, horizon, information availability and retraining schedule of deployment. A metric then translates each forecast error into a business-relevant score.

Validation question

If the deployed model retrains monthly and predicts seven days ahead, should validation fit once and score every later day as a one-step forecast?

Expanding-window backtesting

TrainTrainTrainTrainValid
TrainTrainTrainTrainTrainValid
TrainTrainTrainTrainTrainTrainValid
TrainTrainTrainTrainTrainTrainTrainValid
Training historyNext validation origin

Expanding window

Retains all history. Useful when older observations remain informative and the production model continually accumulates data.

Sliding window

Keeps only the most recent \(w\) observations. Useful when old regimes may be misleading.

Fold isolation: transformations, imputers, scalers, feature selection, lag creation and model fitting must use each fold's training window only.

Error-cost question

Should a model treat an error of 40 rentals as four times or sixteen times as serious as an error of 10?

Training loss defines what “best fit” means

A loss function scores one prediction. Training minimizes the total or average loss. The evaluation metric can match the loss, but it does not have to; choose both deliberately.

Squared-error loss

\[L(y,\widehat y)=(y-\widehat y)^2\]

Large errors receive quadratically more weight. The optimal constant prediction is the mean.

Absolute-error loss

\[L(y,\widehat y)=|y-\widehat y|\]

Every extra unit of error has equal cost. The optimal constant prediction is the median.

Huber loss

\[L_\delta(e)=\begin{cases}\frac12e^2,&|e|\leq\delta\\\delta(|e|-\frac12\delta),&|e|>\delta\end{cases}\]

Quadratic near zero and linear for large errors; less dominated by outliers than squared error.

Quantile or pinball loss

\[L_\tau(e)=\begin{cases}\tau e,&e\geq0\\(\tau-1)e,&e<0\end{cases},\quad e=y-\widehat y\]

Learns a conditional quantile. Use \(\tau>0.5\) when underforecasting is more costly or to construct prediction intervals.

Poisson deviance

\[D=2\left[y\log\!\left(\frac{y}{\widehat y}\right)-(y-\widehat y)\right]\]

Designed for nonnegative count outcomes, with \(\widehat y>0\) and \(y\log(y/\widehat y)=0\) when \(y=0\).

Loss is a modeling decision: squared loss targets the conditional mean; absolute loss targets the median; quantile loss targets the chosen quantile. They can produce different forecasts from the same data.

Evaluation metrics answer different questions

MAE

\[\operatorname{MAE}=\frac1n\sum_{i=1}^{n}|y_i-\widehat y_i|\]

Average absolute miss in rental units. Easy to communicate and relatively robust.

RMSE

\[\operatorname{RMSE}=\sqrt{\frac1n\sum_{i=1}^{n}(y_i-\widehat y_i)^2}\]

Same target units as MAE, but gives large misses more influence.

MAPE

\[\operatorname{MAPE}=\frac{100}{n}\sum_{i=1}^{n}\left|\frac{y_i-\widehat y_i}{y_i}\right|\]

Percentage interpretation, but undefined at zero and unstable near zero.

MASE

\[\operatorname{MASE}=\frac{\frac1n\sum|y_i-\widehat y_i|}{\frac1{T-1}\sum_{t=2}^{T}|y_t-y_{t-1}|}\]

Scales test MAE by a training-set naive error. Below 1 beats that benchmark.

For seasonal data, replace the MASE denominator with the training error from an appropriate seasonal-naive forecast. Always state the scale used.

Relative-error question

Are errors \(10\to20\) and \(100\to110\) equally serious? RMSE says both miss by 10. Would a relative-demand metric agree?

RMSLE: the Kaggle competition metric

\[\operatorname{RMSLE}=\sqrt{\frac1n\sum_{i=1}^{n}\left[\log(1+\widehat y_i)-\log(1+y_i)\right]^2}\]

The \(+1\) allows zero counts. Predictions must be nonnegative. RMSLE measures distance after logarithmic compression, so proportional errors matter more than equal absolute errors.

Actual \(y\)Forecast \(\widehat y\)Absolute errorAbsolute log error
102010\(\left|\log(21)-\log(11)\right|\approx0.647\)
10011010\(\left|\log(111)-\log(101)\right|\approx0.094\)

Both forecasts miss by 10 bikes, but doubling a small demand level is much more serious on the log scale than moving from 100 to 110.

Why it helps

Reduces domination by a few extremely busy hours, works naturally with right-skewed counts, and emphasizes relative accuracy.

What to watch

It can hide large absolute capacity errors at busy hours. A production team may still need MAE, peak-hour error and service-level costs.

Training on \(\log(1+y)\): minimizing MSE on the transformed target resembles RMSLE optimization, but exponentiating a mean log forecast does not automatically recover the mean count. Back-transformation bias and nonnegative predictions need attention.

Uncertainty question

Should a 14-day-ahead interval usually be as narrow as tomorrow's interval?

A point forecast is incomplete

\[\text{point forecast}\quad\widehat y_{t+h\mid t}\qquad\text{and interval}\quad[L_{t+h},U_{t+h}]\]

A nominal 90% prediction interval should contain roughly 90% of future outcomes under conditions represented by evaluation. Check empirical coverage and average width across backtest origins.

Fair comparison: evaluate every candidate on the same origins, horizons, information, metrics and baseline. Also report scores by horizon and by operational segment such as peak hour versus off-peak.

Previous: SARIMA PreNext: ML workflow