Guided tour · start here
The site in three walkthroughs
Three short screen recordings of the main journeys, about 3 minutes in all: browsing the 40 notebooks, asking whether a price model beats yesterday's close, and putting the 2022 spam model against a language model. Each step is listed next to its video as a transcript; press a step to jump to it. Below them are screenshots of every key feature.
Walkthrough 1 of 3 · 46 s
Browse the 40 reps
From the landing page's habit grid to the full log: filter the 40 notebooks by theme and by how far they ran in 2022, then open one to see what it printed and the audit that re-checks it.
No sound; the caption banner is part of the recording. Download the MP4 or the captions (WebVTT).
Steps and transcript
- Data and settings
- Nothing to enter: the run log is generated by scripts/runlog.py from the saved 2022 notebook outputs, without re-running them. Filters: theme Time series, then status Ran. Card opened: 01, Apple price with an LSTM.
- Try it yourself
- Home The 40, filtered Notebook 01
Walkthrough 2 of 3 · 51 s
Can the model beat yesterday?
The stock-forecast audit: five 2022 price models against the flat forecast (tomorrow's close equals today's), first on the shuffled split the notebooks used, then trained on the past and tested on the future, with a Diebold-Mariano test on the same days and the LSTM re-trained with five seeds.
No sound; the caption banner is part of the recording. Download the MP4 or the captions (WebVTT).
Steps and transcript
- Data and settings
- Daily Apple and Netflix prices from Yahoo Finance (yfinance); the last 20% of days are held out in order. LSTM seed 42 with TensorFlow's deterministic ops, plus seeds 1 to 4. Moving-block bootstrap: seed 2026, 10,000 resamples. Diebold-Mariano with the Harvey, Leybourne and Newbold correction.
- Try it yourself
- Audits Stock forecasts Paired tests
Walkthrough 3 of 3 · 1 min 8 s
Spam: classical vs LLM
The one audit where the 2022 model holds up, then the bring-your-own-key harness that puts a language model against it on the same held-out messages. No key is entered: the AI settings dialog is only opened, and the run uses the site's own mock, a keyword rule labelled as such everywhere.
No sound; the caption banner is part of the recording. Download the MP4 or the captions (WebVTT).
Steps and transcript
- Data and settings
- SMS Spam Collection (UCI, CC BY 4.0), the 2022 split with seed 42. The evaluation set is 400 held-out messages drawn with seed 2026; the run uses the first 100 (18 spam), 10 per call. Paired bootstrap: seed 2026, 10,000 resamples. Mocked run for illustration: no API key, no provider call.
- Try it yourself
- SMS spam LLM against naive Bayes AI audit log
Screenshots
Desktop shots at 1440 × 900 in light mode (the landing page in dark mode too) and three phone shots. Select one to enlarge it, then use the arrow keys to move through them. The shots marked “mocked run for illustration” show the site's keyword-rule mock in the LLM's place: no key was entered and no provider was called.
Landing page. The habit grid of 40 notebooks, the run counts and links into the in-browser tools. Landing page, dark mode. The same page in dark mode: chalk on slate. The 40, filtered. Time-series notebooks that ran in 2022: theme and status chips, search and live counts. One notebook's run record. What notebook 01 printed in 2022, read from its saved outputs, and what I'd change. Eight audits. Verdict counts and one card per audit; the interactive ones are marked. LSTM against yesterday's close. The 2022 shuffled split against a time-ordered split, with a no-training baseline. Is the gap bigger than the noise? Block-bootstrap intervals and Diebold-Mariano tests against the flat forecast. Five seeds, not one. LSTM error for each training seed: the shuffled-split verdict depends on the seed. Spam model with intervals. Wilson and bootstrap intervals, and McNemar's test against 'always ham'. Try the spam model. Naive Bayes ported to TypeScript: the verdict and the words that moved it. Bring your own key. Anthropic by default or OpenAI; the key stays in this browser and never reaches this site. LLM against naive Bayes (mocked run). Mocked run for illustration: a labelled keyword rule, scored with intervals and McNemar's test. AI audit log (mock entries). Mocked run for illustration: each call with its input, output and human decision; JSON or CSV export. Accuracy paradox explorer. Move the fraud rate and see how much a 99.97% accuracy leaves open. Methods. Provenance, uncertainty, limitations, data and model cards and the AI use statement. Decision record DR-004. The BYOK LLM comparison: the decision first, then options, what happened and changes.
On a phone (390 px wide)
Mobile: landing. The landing page at 390 px. Mobile: the 40. Search and filter chips on a phone. Mobile: LLM harness (mocked run). Mocked run for illustration: the score table stacks on a phone.
How these were made
Every frame comes from a script, not a screen-capture session. A Playwright tour (web/e2e/showcase.spec.ts) drives the site in Google Chrome at 1280 × 800, adds the caption banner and the cursor highlight, and records the video. It also checks what it shows: the four time-series notebooks that ran, Apple's LSTM error on shuffled and time-ordered days, the Diebold-Mariano statistic and interval, the seed table, the spam model's interval and McNemar's test, and the mock run's scores. A broken feature fails the tour rather than producing a misleading video.
The numbers are deterministic: the run log is read from the saved 2022 outputs, the audit results are committed JSON written by seeded scripts, and the mock is a fixed keyword rule on a fixed evaluation set, so a rerun records the same values. No API key is entered anywhere. The AI settings dialog is only opened, and the LLM comparison is run with the mock, which is labelled “Mock, not AI” on the page and in the audit log.