Decision record · DR-003
Re-run audits only on licence-clear data, publish derived numbers, and make one attributed exception for the SMS corpus
- Status
- Accepted
- Date
- 2026-10
- Applies to
- scripts/, web/src/data/, web/public/data/, NOTICE, /about#data, /methods#data-cards
Decision in one line
An audit re-runs a notebook only when its data is licence-clear and downloadable without an account; the site publishes metrics, counts and error curves, never raw records, with one exception: a masked, attributed sample of 400 messages from the CC BY 4.0 SMS Spam Collection, because comparing an LLM with naive Bayes needs the text.
Context
The 40 notebooks read a wide mix of data: files from the tutorial author's GitHub repository with no stated licence, Kaggle datasets that need a login, tweet collections with user names and profile fields, mental-health posts, the MovieTweetings ratings and Yahoo Finance prices. The repository is private, but the site is public, and anything it serves is published.
The upgrade added an optional evaluation in which a language model labels SMS messages next to the 2022 naive Bayes model. A language model cannot label a message it is not shown, so for the first time the site needed to serve raw text.
Decision
- Audits re-run code only on data that is licence-clear and downloadable without an account: the SMS Spam Collection, Iris and Automobile (UCI, CC BY 4.0), Water Quality (Kaggle, CC0), the Advertising data from An Introduction to Statistical Learning, Yahoo Finance prices via yfinance, and the price files bundled with the notebooks.
- Downloads go to the git-ignored
scripts/.cache/and are never committed. - The site publishes derived numbers only: metrics, counts, histograms of scores, daily error series and, for the spam playground, per-class token counts with every token of five or more digits removed.
- Fraud (02) and the classifier comparison (27) are worked out from the 2022 outputs, because PaySim needs a Kaggle login and the social network ads table has no stated licence.
- Tweets, Reddit posts, chat logs and user handles are never processed or served.
- One exception:
web/public/data/spam-eval-set.jsonholds 400 messages drawn with seed 2026 from the 2022 test split of the SMS Spam Collection, with every run of five or more digits replaced by#####. The file names the corpus, its authors, the licence and the changes, and the site repeats that attribution wherever the messages appear.
Options considered
- Publish no raw text, and drop the LLM comparison. The safest option, but it removes the evaluation the upgrade is meant to show.
- Let visitors fetch the corpus from UCI in their own browser. It keeps the text off this site, but the UCI archive is a zip file that the browser cannot fetch across origins, and the sample would depend on a third-party URL staying put.
- Publish a small, masked, attributed sample (chosen).
- Write synthetic messages. No licence question at all, but they would test the model on my idea of spam rather than on the benchmark the 2022 notebook used.
Why
CC BY 4.0 allows redistribution and adaptation with attribution and an indication of changes, and the corpus was built and published as a public research resource. Four hundred messages is enough for an evaluation and small next to the 5,574 the authors publish. Masking long digit runs removes the phone numbers and premium short codes in the spam, the same rule the browser model's vocabulary already follows, and it does not touch what makes a message spam or ham.
What happened
- 47 of the 400 messages had a digit run masked, and masking changed none of naive Bayes's 400 predictions, so the comparison is unaffected.
- Some of the spam is crude, and some personal messages mention first names. The corpus has been public since 2011; the evaluation section warns that the messages are real.
- The sample comes from a public benchmark, so a language model may have seen these exact messages, with their labels, during training. The evaluation card treats that as its first limitation.
- The Water Quality file is read from the tutorial author's repository, which states no licence, because Kaggle needs a login. The script checks that its size matches the 525,187 bytes Kaggle reports for the CC0 original, and the site publishes only metrics from it.
- Notebook 27's audit needed one number the 2022 output never printed: how many training rows a
fully grown tree can get right.
scripts/audit_classifier_scores.pydownloads the social network ads table into the cache, re-creates the 2022 split and publishes only counts (rows per split, class sizes, the training ceiling of 359 of 360, the share of seeds that reproduce each printed score) with the file's checksum. No rows or feature values leave the machine. It moved the implied test accuracy from 32 to 33 of 40.
What I'd change
- Add a contamination check: compare the LLM on these messages with the same model on fresh, hand-written messages of the same kinds.
- Record a checksum of each downloaded file next to the derived numbers for every audit, as the spam, water, sales and car scripts already do.