Decision record · DR-002
Judge every price forecast on later days, against a flat forecast, with a paired test
- Status
- Accepted
- Date
- 2026-10
- Applies to
- scripts/audit_forecast_*.py, web/src/lib/audit-rigour.ts, /audits/forecasts
Decision in one line
Every price model in the log is scored on days after the ones it learnt from, next to a flat "no change" forecast scored on the same days, and the gap is tested with the Diebold-Mariano test on the daily errors rather than read off two averages.
Context
Five notebooks predict share prices: LSTMs for Apple and Netflix (01, 09), a straight line five
days ahead (05) and AutoTS five-day forecasts for Tata Motors and Apple (28, 32). Notebooks 01,
05 and 09 shuffled the days with train_test_split before fitting, so every test day sat
between training days on both sides. Notebooks 01 and 09 also fed in the same day's high and
low, which bracket the close they were asked to predict. Notebook 28 printed a forecast that was
never checked, and 32 never finished. None compared the model with the simplest forecast there
is: tomorrow's price equals today's.
Decision
- Keep the 2022 shuffled split, to score the recipe as it was run, and add a time-ordered split: the first 80% of days for training, the last 20% for testing, in order.
- Score every model against a flat forecast on exactly the same days: yesterday's close (01, 09), today's price five days ahead (05) and the last close held for five days (28, 32). For 01 and 09 also show (high + low) / 2, which needs no training and exposes the leak.
- For AutoTS, repeat the 2022 call on four consecutive five-day windows at the end of each bundled year.
- Report the mean absolute percentage error (MAPE) with a 95% moving-block bootstrap interval, and test model against flat forecast with the Diebold-Mariano test (Harvey, Leybourne and Newbold correction), absolute percentage error as the loss and the forecast horizon as h. For AutoTS, with four forecast origins, compare window by window with an exact sign test.
- Train the LSTMs with five seeds, because one run proved not to be repeatable (DR-005).
Options considered
- Keep the shuffled split. It answers a question nobody asks of a forecast.
- One time-ordered holdout per model (chosen for 01, 05 and 09). Cheap, and it matches how a forecast is used.
- A full rolling-origin backtest for every model. The best design, but the LSTMs train with batch size 1 for 30 epochs, so hundreds of origins would take days of CPU.
- Rolling windows for AutoTS (chosen, four windows). Each AutoTS fit takes several minutes, which limited the number of windows.
Why
A forecast is used on days after the data it learnt from, so that is where it should be tested. For daily prices the flat forecast is famously hard to beat, which makes it the honest bar. Comparing two average errors ignores that the errors of nearby days move together, and that five-day forecasts overlap; the Diebold-Mariano test accounts for both, and testing the daily differences pairs the comparison day by day.
What happened
- On the time-ordered split the Apple LSTM misses by 38.7% on average against 1.7% for yesterday's close (Diebold-Mariano statistic 53.7 over 690 days). The network never saw a close as high as most of the test days, and a model trained on raw prices cannot extrapolate.
- On the shuffled split the Apple LSTM beats yesterday's close with seed 42 (1.00% against 1.29%), but averaging the same day's high and low does better still (0.59%) with no training at all, and with the four other seeds the LSTM scores 1.30% to 2.27%, no better than yesterday's close.
- The straight line in notebook 05 reproduces its 2022 R² exactly, then loses to the flat forecast on later days: 0.50% against 0.36% (Diebold-Mariano 2.77, p = 0.008, h = 5).
- AutoTS lost to the flat forecast in all four windows for both shares, but four windows are not much evidence: the exact sign test cannot go below p = 0.125 with four windows. The 20 days are four forecasts of one to five days ahead, not 20 comparable days, so the Diebold-Mariano test is not applied to them.
- Netflix did not repeat. Re-running the LSTM script in October 2026, with the same code, data, seed and TensorFlow build, reproduced Apple to every published digit but moved Netflix's MAPE from 8.48% to 4.09% (shuffled) and from 18.22% to 14.57% (time-ordered). The conclusion held in both runs, but a single run is not a result, so every fit was repeated with four more seeds. On the time-ordered split all five seeds lose heavily: Apple 35.9% to 43.3%, Netflix 10.9% to 16.4%, against 1.7% and 2.2% for the flat forecast.
What I'd change
- Use a rolling origin for every model, with enough origins to put an interval on the difference in skill.
- Model returns, not raw prices, and scale the inputs, so that the LSTM is not asked to extrapolate.
- Report a skill score against the flat forecast as the headline, with MAPE second.