Backtesting forecasts in practice: Define forecast measure, unit, frequency and lead time before testing; Use only data available at each issue date to avoid information leakage; Compare errors across origins using consistent units and outcome rules
Image: Business Insight Stack

Data Modelling

Part of Analytical forecasting

Backtesting a forecast using earlier periods

Recreate historical forecast issue dates, prevent future-data leakage, compare errors at the right horizon and interpret the result’s limits.

Backtest a forecast by recreating several earlier issue dates. At each date, limit the method to information available then, predict the required future period and compare the prediction with what happened. The exercise is useful when its cutoff, horizon and outcome definition resemble the decision the forecast now serves.

Define the forecast task first

Record the measure, unit, frequency, source cutoff and lead time. A four-week staffing forecast should be assessed across four weeks, even if the method also performs well one week ahead. Decide whether success means accurate individual weeks, the four-week total or both.

Define the outcome consistently. If “completed cases” excludes reopened cases today but included them in older reports, settle which rule the backtest uses before interpreting errors. Mark missing deliveries and major definition changes; an absent period should not be treated as an ordinary zero without evidence.

Move the issue date forward

This schedule illustrates the method; it reports no test result:

Issue dateInformation allowedOutcomes later compared
End of JuneRecords available by that cutoffJuly weeks 1–4
End of JulyRecords available by that cutoffAugust weeks 1–4
End of AugustRecords available by that cutoffSeptember weeks 1–4

Fit or update each candidate as it would have been updated in practice. Repeat the same steps for the baseline. If the process would retrain monthly, retrain monthly in the exercise. Preserve the forecast produced at each origin, along with the data and settings used, before matching it to the later outcome.

External predictors create a particular leakage risk. A calendar field known in June may be usable for July. July’s actual sales, a later-revised demand estimate or a label derived from the final outcome was not available at the June issue date. Using those values measures an ex-post exercise, not an achievable advance forecast.

Compare the relevant errors

For each issue date and horizon, calculate actual minus forecast using the same unit and population. Mean absolute error summarises the size of misses. Root mean squared error gives large misses more weight.

Inspect positive and negative errors for persistent under- or over-forecasting. Avoid relying on percentage error when actual values are zero or close to zero.

Group results by lead time. A method can work one week ahead and poorly four weeks ahead. Compare the candidate and baseline over the same origins, and examine consequential misses as well as average error.

If the method issued prediction intervals, count how often outcomes fell within them at the relevant horizon and whether misses cluster on one side. Historical coverage is a diagnostic of the method and conditions tested, not a promise of future coverage.

Forecast Accuracy Metrics Comparison

  • Mean Absolute Error (MAE)Sum of absolute differences between actual and forecast, averaged across all issue dates
  • Root Mean Squared Error (RMSE)Squared differences weighted more heavily; penalises large errors
  • Prediction Interval CoveragePercentage of outcomes falling within predicted intervals at each horizon

Interpret the result within its limits

A backtest from stable past conditions may say little about a new process or market shock. Few origins make apparent differences between methods fragile. Document excluded periods, source revisions and settings chosen after seeing results. A historical exercise is not a production result.

Where enough data remain, keep a later period untouched while choosing among candidates, then assess the chosen method on it. Once forecasting is in use, archive each issued forecast, assumptions and cutoff. Those records make later comparisons more faithful to what decision makers actually knew.

More from Data Modelling

Data Modelling

BI data modelling

Define reporting grain, facts, dimensions, dates, measures and changing attributes in a BI model that produces explainable figures.