Decision record · DR-005
Put a stated interval, a paired test and a seed on every audit number
- Status
- Accepted
- Date
- 2026-10
- Applies to
- web/src/lib/stats, web/src/lib/audit-rigour.ts, scripts/common.py, scripts/stats_reference.*, every /audits page
Decision in one line
Every audit number now carries a 95% interval chosen for what it measures (Wilson for proportions, a seeded bootstrap otherwise, moving blocks for daily errors, a corrected resampled t for cross-validation means), every comparison is paired on the same items (McNemar for classifiers, Diebold-Mariano for forecasts), and every seed is printed next to the result.
Context
The first version of the audits already made the main point of each notebook, often with a spread over random splits. But most numbers were point estimates, comparisons were read off two averages, and the reader could not tell how much of a difference was noise. Re-running the LSTM script then showed that even "the same run" did not always give the same number.
Decision
- Proportions (accuracy, precision, recall, a baseline's hit rate) get Wilson score intervals, which stay inside 0 to 1 and behave well for small samples and rates near 100%.
- Other statistics on items (F1, R², mean absolute error, ROC AUC, paired differences) get percentile bootstrap intervals over items, seed 2026, 10,000 resamples (2,000 where every resample refits an expensive statistic, which the page states). F1 is resampled within each true class, so every resample keeps the observed number of positives, and it gets no interval with fewer than 10 positives or when a model makes no errors.
- Daily gaps between two forecasts get the same moving-block bootstrap as the errors, so the effect size next to each Diebold-Mariano test has its own interval.
- Daily forecast errors get a moving-block bootstrap, with blocks of at least the forecast horizon and about n^(1/3) days, so the interval keeps the autocorrelation of the errors.
- Cross-validation means get the corrected resampled t interval of Nadeau and Bengio (2003), because folds share most of their training rows and the naive standard error is too small. The spread across folds is shown as a spread, never as a confidence interval.
- Paired comparisons: McNemar's exact test for two classifiers on the same items, with the conditional odds ratio and the accuracy difference as effect sizes; the Diebold-Mariano test with the Harvey, Leybourne and Newbold correction for a forecast against the flat forecast on the same days; the corrected resampled t on per-fold differences for two models on the same folds; an exact sign test across AutoTS windows, whose 20 days are only four forecast origins.
- Seeds for every split, fold assignment, bootstrap and model are printed on each page. The LSTMs are trained with five seeds and reported as a spread.
- The TypeScript helpers are tested against SciPy, statsmodels and R's
forecast::dm.test, on synthetic inputs and on the exact series the pages show.
Options considered
- Point estimates with the spread across splits (the first version). Honest about variation, but not a statement of uncertainty about one model.
- Normal-approximation intervals everywhere. Simple, but wrong near 100% accuracy and for statistics like F1 and R².
- Bootstrap everything. One method is tidy, but the plain bootstrap ignores the autocorrelation of daily errors and the overlap of cross-validation folds.
- An interval chosen for each kind of number (chosen).
Why
Each interval answers the question its number raises. A paired test uses the fact that both methods saw the same items, which is far more sensitive than comparing two intervals, and an effect size next to every p-value keeps the size of a difference in view. Printing seeds makes every number reproducible from the scripts, and testing the helpers against the standard implementations means the site's statistics are checked, not trusted.
What happened
- Several comparisons that looked like differences are not distinguishable from zero: k = 15 neighbours against k = 1 on Iris (+1.3 points, interval −2.0 to +4.6), dropping newspaper from the sales model (+0.001 R²), and AutoTS against the flat forecast, which lost all four windows for both shares (exact sign test p = 0.125, the lowest four windows allow).
- Logistic regression on the water data cannot be told apart from "always unsafe" (McNemar p = 0.51), which the accuracy figures alone only hinted at.
- The tree beats linear regression on car prices by 0.085 R² across folds, with an interval of 0.001 to 0.169: a real but loosely measured advantage.
- Re-running the LSTM script with the same seed reproduced Apple exactly and Netflix not at all (DR-002), which is why the LSTMs are now reported across five seeds. The seeds matter: on the shuffled split, the Apple LSTM beats yesterday's close with the seed the audit first used and with none of the other four.
- Review before release caught two places where the first choice was wrong. A plain bootstrap of F1 on the LLM comparison's smallest sample (25 messages, 2 of them spam) printed "0.000 to 1.000" for a model that made no mistakes, because about one resample in eight had no spam at all; F1 is now resampled within each class and not given an interval at that size. And a Diebold-Mariano test on AutoTS's 20 stacked days gave p = 0.06, which read as near-significant evidence from what are really four forecasts; the sign test replaced it.
- The site runs many tests and does not correct for multiple comparisons. Each p-value is descriptive for its own question, and the methods page says so.
What I'd change
- Check the coverage of the moving-block intervals by simulation on series with the same autocorrelation, and pick the block length from the data rather than a rule of thumb.
- Pre-register which comparisons are confirmatory and correct those for multiplicity.
- Replace the four AutoTS windows with a rolling origin long enough to give the Diebold-Mariano test real power.