Decision record · DR-001
Present the 40 tutorial notebooks as an honest run log with audits, not as a portfolio of projects
- Status
- Accepted
- Date
- 2026-10
- Applies to
- original/, scripts/runlog.py, web/src/data/runlog.json, /projects, /audits
Decision in one line
The site reports what the 2022 notebooks actually did, read by a script from their saved outputs, and adds eight audits that re-check results where a baseline, a time-ordered split or an interval could change the conclusion; it never republishes the tutorials' code or narrative, and it never edits a 2022 number.
Context
In November and December 2022 I worked through 40 of Aman Kharwal's tutorials, one notebook a day, to build fluency with the everyday tools. The notebooks follow the tutorials closely, so their code and much of their narrative are the tutorial author's work, and several hold third-party data I am not allowed to share. My 2022 README presented them as a list of projects. Coming back in 2026 I found that some had never finished, and that many headline numbers were measured in ways that flatter the model: shuffled time series, scores on training rows, no baseline, 30-flower test sets.
I wanted to show this body of work without passing the tutorials off as mine and without hiding the weak parts, because the weak parts are where the learning is.
Decision
- The notebooks move unchanged into
original/, and the repository stays private. scripts/runlog.pyreads every notebook's saved outputs, without re-executing anything, and records which cells ran, which failed and why, and the numbers each notebook printed. It asserts the expected status of every notebook, and CI checks that the committed run log still reproduces fromoriginal/.- Each project page describes the exercise in my own words and links to the tutorial it follows. No notebook code or tutorial text appears on the site.
- Eight audits re-check notebooks where a little more rigour could change the conclusion. Six re-run the method on licence-clear data; two are worked out from the numbers the 2022 notebook printed, because the data needs an account or has no stated licence (DR-003).
- Each audit carries a verdict: the conclusion flips, it is weaker than it looked, it holds, or the 2022 output cannot tell. The 2022 numbers are shown as printed; the audits add analysis around them and never replace them.
Options considered
- Publish the notebooks as portfolio projects. Quick, but the code is the tutorial author's, several datasets cannot be redistributed, and the headline numbers would mislead.
- Re-run all 40 notebooks today and report the new outputs. It would rewrite the record of what happened in 2022, half the environments no longer build, and most datasets cannot be shared.
- Leave the notebooks out of my portfolio. Safe, but it throws away the most useful lessons in them.
- An honest run log plus targeted audits (chosen).
Why
The run log is a fact I can check, not a claim: a script extracts it and CI reproduces it. The audits are where the value is, because they show the habits I would want a reviewer to see: writing down the baseline first, splitting by time when the data has an order, and putting an interval on a small test set. Keeping the tutorials' work off the site respects their author and keeps the repository private for a clear reason.
What happened
- Of the 40 notebooks, 33 ran from top to bottom, 3 stopped part-way, 3 failed at the start and 1 was never written. The old README said nothing of the seven that did not finish.
- Of the eight audits, three conclusions flip, three are weaker than they looked, one cannot be judged from the 2022 output and one holds: the spam model in notebook 37.
- That count is not a measure of the 40 notebooks. I chose the eight audits partly because they looked suspicious, so "three of eight flip" says something about my selection, not about the tutorials.
- Several audits run on a different copy of the data than the notebook read (the UCI files rather than re-uploads), so their numbers follow the same recipe on slightly different rows. Each audit page says so where it applies.
What I'd change
- Write each audit's question and analysis plan down before running it, so the verdicts could not be shaped by what I found.
- Audit a random sample of notebooks as well as the suspicious ones, to say something about the set as a whole.
- Check the tutorial links in CI. Some of the tutorial site's pages refuse automated requests, so this needs a polite, rate-limited check rather than a plain link checker.