Methods
How the re-check works, and what it cannot tell you
Where the data comes from, how each number is tested, how uncertain it is, the decisions behind the design and how the optional AI feature is governed.
This page sets out where every number on the site comes from, how each one is tested, how uncertain it is, what I assumed, and what the work cannot tell you. The decisions behind the design are recorded separately, in the decision records.
Provenance
The site has three layers, and each has its own source of truth.
| Layer | Source | How it is produced | What is published |
|---|---|---|---|
| The run log | The saved outputs of my 40 notebooks from late 2022, kept unchanged in original/ | scripts/runlog.py reads the outputs without re-running anything; CI checks the log still reproduces | Statuses, cell counts, error names and the numbers each notebook printed |
| Six recomputed audits | Licence-clear public data, downloaded by the scripts into a git-ignored cache | One uv script per audit re-runs the 2022 recipe and adds the missing baseline, split or interval | Metrics, counts, score histograms, daily error series |
| Two explained audits | The numbers printed by notebooks 02 and 27 | Worked out from those outputs and checked in code; for 27 a local script re-creates the 2022 split to count what a tree can learn | The arithmetic, counts and the intervals |
Each dataset has a card below, with its licence, what the audit uses and what the site publishes. The tutorials themselves belong to Aman Kharwal; the site links to each one and does not republish their code or text (DR-001).
Method
- Reproduce first. Each audit script starts from the 2022 recipe: the same columns, model, split fraction and random seed. Where the 2022 number can be reproduced it is: the linear forecast script (05) refuses to continue unless it matches the 2022 R² exactly, the sales script (03) records that it does, and the LSTM script checks row counts and closing prices against what the 2022 notebooks printed.
- Add the missing comparison. A majority-class baseline for every classifier, a flat forecast for every price model, the average price for the car model.
- Test the way the model would be used. Time-ordered splits for anything with dates (DR-002), repeated cross-validation where a single small split decides the score, and out-of-fold predictions for thresholds.
- Put an interval and a paired test on every number, as described next (DR-005).
- Publish derived numbers only (DR-003).
Statistics that need rows the site does not publish (ROC AUC, R² on one split) are bootstrapped
in the Python scripts. Everything that can be computed from the published counts and series is
computed on the site by the TypeScript code in web/src/lib/stats, at build time for the audit
pages and in the browser for the LLM comparison.
Uncertainty
Every interval on the site is 95%. Which one depends on what is being measured.
| Measured | Interval or test | Why this one |
|---|---|---|
| A proportion: accuracy, precision, recall, a baseline's hit rate | Wilson score interval | Stays inside 0 to 1 and keeps close to its stated coverage for small samples and rates near 100% |
| Another statistic over items: F1, R², mean absolute error, ROC AUC, a paired difference | Percentile bootstrap over items, seed 2026, 10,000 resamples (2,000 where each resample is expensive) | Makes no distributional assumption about the statistic |
| F1 and differences in F1 | The same bootstrap, resampled within each true class, so every resample keeps the observed number of positives; no interval with fewer than 10 positives or when a model makes no errors | A resample with no positives has no F1, and a model with no errors gives a zero-width interval; the Wilson bounds on precision and recall say more there |
| A mean of daily forecast errors, or of the daily gap between a model and the flat forecast | Moving-block bootstrap, blocks of at least the forecast horizon and about n^(1/3) days | Keeps the autocorrelation of errors on nearby days |
| A cross-validation mean | Corrected resampled t (Nadeau and Bengio, 2003) | Folds share most of their training rows, so the naive standard error is too small |
| Two classifiers on the same items | McNemar's exact test, with the conditional odds ratio and the accuracy difference | Only the items where they disagree carry information |
| A forecast against the flat forecast on the same days | Diebold-Mariano test with the Harvey, Leybourne and Newbold correction; absolute percentage error as the loss; h = the forecast horizon | Allows for autocorrelated and overlapping forecast errors |
| Two models on the same cross-validation folds | Corrected resampled t on the per-fold differences | Pairs the comparison fold by fold |
| AutoTS against the flat forecast over four windows | Exact sign test, with the lowest and highest window in place of an interval | The 20 days are four forecasts of one to five days ahead, so there are four origins, not 20 comparable days; a Diebold-Mariano test or a bootstrap on the stacked days would claim more than that |
Seeds are printed on every audit page: the 2022 split seeds, the fold seeds, the bootstrap seed
(2026) and, for the LSTMs, the five training seeds. The TypeScript helpers are tested against
SciPy and statsmodels (scripts/stats_reference.py) and against R's forecast::dm.test
(scripts/stats_reference.R), both on synthetic inputs and on the exact error series the
forecast page shows.
A p-value on this site describes one comparison. The site runs many of them and does not correct for multiple comparisons, so read the p-values alongside the effect sizes and intervals, not as a scoreboard.
Assumptions
- Test rows are treated as independent draws, which the Wilson intervals and the item bootstrap need. For daily prices this is false, which is why the forecast page uses moving blocks and the Diebold-Mariano test instead.
- The Diebold-Mariano test assumes the daily loss difference is stationary over the test period. Prices in 2020 to 2022 moved a lot, so the forecast p-values are a guide, not a guarantee.
- Where an audit uses the UCI copy of a dataset rather than the re-upload the notebook read, I assume the few differing rows do not change the conclusion. Each page says which copy it used.
- The fraud audit takes PaySim's class count (8,213 frauds in 6,362,620 rows) from the figure widely reported for the file, because the data needs a login.
- The classifier-scores audit backs the tree's test accuracy out of a score on all 400 rows. That needs the tree's training accuracy, which the notebook never printed; a local re-creation of the 2022 split puts it at 359 of 360 (one conflicting pair) for every seed tried. The implied 33 of 40 is only as good as that re-creation, which reproduces all four printed scores.
Limitations
- Small test sets stay small. An interval makes 30 flowers or 39 cars honest; it does not make them informative.
- Selection. I chose the eight audits partly because they looked suspicious, so the count of conclusions that flip describes my choice, not the 40 notebooks.
- The LSTMs are fragile. Re-running the same script with the same seed reproduced Apple exactly and Netflix not at all, so the LSTM numbers are reported across five seeds. On the shuffled split, whether the LSTM beats yesterday's close depends on the seed.
- Four AutoTS windows cannot establish anything at the 5% level; the direction is consistent, the evidence is thin.
- A majority-class baseline is a low bar. Beating it is necessary, not impressive, and the McNemar tests against it mostly confirm that a model learnt something.
- No multiplicity correction, as above.
What I'd change
- Write each audit's question and analysis plan down before running it.
- Use a rolling-origin backtest for every forecast, with enough origins to put an interval on the difference in skill.
- Check the coverage of the moving-block intervals by simulation.
- Audit a random sample of the 40 notebooks as well as the suspicious ones.
Data cards
One card for each dataset an audit uses, including the two it does not re-run or republish. Licence choices are explained in DR-003. Raw downloads stay in the scripts' git-ignored cache; the site publishes only what each card lists under "Published".
SMS Spam Collection
- Audit: spam detection (37), and the LLM comparison on the same page.
- Source and licence: Almeida, Gómez Hidalgo and Yamakami (2011), UCI Machine Learning Repository, CC BY 4.0. 5,574 English text messages, 747 of them spam.
- Collection: spam from a UK complaints forum (Grumbletext), personal messages from the NUS
SMS Corpus (Singapore students) and other public sources, all before 2011. Abbreviations,
Singlish and placeholder tokens such as
<#>are common. - Used for: the 2022 pipeline's held-out score (1,840 test messages, split seed 42), 10-fold cross-validation (seed 0), the browser model, and the evaluation set for the LLM comparison (400 test messages, seed 2026).
- Published: metrics and score histograms; the browser model's per-class token counts with
every token of five or more digits removed; 400 messages with five-or-more-digit runs masked
as
#####, attributed to the authors. - Caveats: dated and regional; labels come from the source, not from a fresh review, and a few are debatable; some spam is crude. The corpus is a well-known benchmark, so language models may have seen it.
Water Quality
- Audit: water quality (24).
- Source and licence: "Water Quality" by Aditya Kadiwal on Kaggle, CC0. 3,276 water samples with nine measurements and a potability label. Read from the copy the 2022 notebook used, checked against the size Kaggle reports, because Kaggle needs a login.
- Used for: 5-fold cross-validation (seed 0) of logistic regression and a random forest (seed 0) on the 2,011 complete rows, and the forest on all rows with median imputation.
- Published: accuracy, confusion counts, ROC AUC with bootstrap intervals, and histograms of out-of-fold probabilities.
- Caveats: the dataset's provenance is thin: no sampling frame, locations or dates are given, so the results describe this table, not drinking water in any particular place.
Iris
- Audit: Iris (23, 15).
- Source and licence: Fisher (1936) via the UCI Machine Learning Repository, CC BY 4.0, using the copy bundled with scikit-learn. 150 flowers, three species. 148 rows match the copy the 2022 notebooks read; the other two are the known corrections in the UCI file.
- Used for: the 2022 split (seed 0, 30 test flowers), 1,000 other splits (seeds 0 to 999) and 5-fold cross-validation repeated 50 times (seed 0).
- Published: counts of correct predictions, intervals and cross-validated accuracies.
- Caveats: a famously easy, tiny dataset. It is useful here precisely because it shows how little 30 test rows can say.
Automobile (1985 Auto Imports Database)
- Audit: car price (36).
- Source and licence: Jeffrey Schlimmer via the UCI Machine Learning Repository, CC BY 4.0. 205 cars; 195 have every column the 2022 model used.
- Used for: the 2022 split (seed 42, 39 test cars) with 50 tree seeds, and 5-fold cross-validation repeated 20 times (seed 0) for a tree and a linear model.
- Published: R², mean absolute error in dollars and their intervals.
- Caveats: 1985 list prices for imported cars in the United States; a small table where a handful of luxury models dominate the errors.
Advertising
- Audit: sales from advertising (03, 14).
- Source: the Advertising data from An Introduction to Statistical Learning (James, Witten, Hastie and Tibshirani), downloaded from the book's website. 200 markets with TV, radio and newspaper budgets and sales.
- Used for: the 2022 split (seed 42), 1,000 other splits and 5-fold cross-validation repeated 20 times (seed 0). The re-upload the notebook read differs from the book in 59 sales values; it is used only to confirm the 2022 R² reproduces exactly.
- Published: R² values, their distribution and intervals, and the model's coefficients.
- Caveats: a teaching dataset with no stated licence beyond the book's site, so only derived numbers are shown.
Share prices
- Audits: stock forecasts (01, 05, 09, 28, 32).
- Sources: daily Apple and Netflix prices from Yahoo Finance via yfinance, the same request as in 2022, with Netflix's 2025 split undone so the scale matches (01, 09); the price files bundled with notebooks 05, 28 and 32, which came from the tutorial author's repository with no stated licence.
- Used for: shuffled and time-ordered splits, four rolling five-day windows, and the week after notebook 32's data ends.
- Published: error metrics and daily absolute percentage errors only, never price tables.
- Caveats: Yahoo's data is provided under Yahoo's terms. Tata Motors' American depositary receipt was delisted in 2023, so the 2022 forecast for it can no longer be checked.
PaySim (not re-run)
- Audit: fraud (02), worked out from the numbers the notebook printed.
- Source and licence: synthetic mobile-money transactions by Edgar Lopez-Rojas, Kaggle, CC BY-SA 4.0. Downloading it needs a Kaggle login, so the audit does not.
- Caveats: the class count used (8,213 frauds in 6,362,620 rows) is the figure widely reported for the file. The dataset's own notes say the balance columns record the outcome of cancelled frauds, which the 2022 model used as features.
Social network ads (re-created locally, counts only)
- Audit: four classifiers (27), worked out from the four scores the notebook printed.
- Source: a 400-row table of age, estimated salary and purchase from the tutorial author's repository, with no stated licence.
- Used for: a local check only.
scripts/audit_classifier_scores.pydownloads the table into the git-ignored cache, re-creates the 2022 split (seed 42, 360 training and 40 test rows) in a 2022-era scikit-learn, confirms all four printed scores and counts how many training rows a fully grown tree can get right (359 of 360, because of one conflicting pair). The tree's own seed was never set, so it tries 200 seeds. - Published: counts only: rows per split, class sizes, the training ceiling and the share of seeds that reproduce each printed score. No rows or feature values.
- Caveats: with 40 test rows the implied test accuracy is a wide interval, and it rests on the re-created split matching the 2022 one, which the four reproduced scores support but cannot prove.
Model card: the spam model in your browser
The spam playground on /audits/spam runs a multinomial naive Bayes model in the visitor's
browser. This card describes that model. It is the only model the site runs for a visitor; the
LLM comparison uses models the visitor chooses, described in the
evaluation card.
Model
- Type: multinomial naive Bayes (scikit-learn
MultinomialNB, alpha = 1) on word counts from scikit-learn's defaultCountVectorizer(lowercase, tokens of two or more word characters). - Implementation: a TypeScript port of the tokeniser and
predict_proba(web/src/lib/naive-bayes.ts), driven by per-class token counts inweb/public/data/spam-model.json. A parity test checks it against scikit-learn's own probabilities on twelve hand-written messages. - Output: the probability that a message is spam, the verdict at 50%, and each known word's contribution in log-odds.
Intended use
- To let a visitor see how a classical text classifier reaches a verdict, word by word.
- Not intended for filtering anyone's real messages, for moderation, or for any decision about a person.
Training data
- The SMS Spam Collection (UCI, CC BY 4.0): 5,574 English text messages collected before 2011, 747 of them spam. See the data card.
- The browser model is fitted on all 5,574 messages. Tokens containing five or more digits in a row (phone numbers and premium short codes) are removed from the published vocabulary before fitting, so the published counts carry no numbers that could identify anyone.
Evaluation
The browser model is trained on every message, so it has no held-out set of its own. Its evaluation is the 2022 pipeline's, which uses the same features and model: trained on two thirds of the messages (split seed 42) and tested on the other 1,840.
| Measure | Estimate | 95% interval | Method |
|---|---|---|---|
| Accuracy | 98.4% | 97.7% to 98.9% | Wilson, 1,811 of 1,840 |
| Spam precision | 93.1% | 89.4% to 95.6% | Wilson, 244 of 262 |
| Spam recall | 95.7% | 92.4% to 97.6% | Wilson, 244 of 255 |
| Spam F1 | 0.944 | 0.923 to 0.963 | Percentile bootstrap within each true class, seed 2026, 10,000 resamples |
| "Always ham" baseline | 86.1% | 84.5% to 87.6% | Wilson, 1,585 of 1,840 |
- Against "always ham" on the same messages, McNemar's exact test gives p < 0.0001: the model alone is right on 244 messages, the baseline alone on 18.
- Stratified 10-fold cross-validation (seed 0, vocabulary learnt inside each fold) gives a mean accuracy of 98.7%, with a corrected resampled t interval of 98.1% to 99.3%.
- 18 of the 1,585 genuine test messages (1.1%) were flagged as spam. For a filter, that is the number to watch.
Known failure modes
- Short or unfamiliar messages. With no known words the verdict falls back to the class balance, so it says "ham".
- New kinds of spam. The vocabulary is from before 2011: no crypto, no delivery-scam links, no modern brand names.
- Spelling games. "Fr33 pr1ze" shares no tokens with "free prize".
- Overconfidence. Naive Bayes treats words as independent, so its probabilities are extreme; read them as a ranking, not as calibrated chances.
- Language and region. Trained on UK and Singapore English; other dialects and languages are out of scope.
Ethical considerations
- A false positive hides a genuine message, which usually costs the recipient more than one extra spam. The threshold explorer on the audit page shows the trade-off.
- The training messages are real texts from a public research corpus. The browser model publishes only aggregate token counts, with long numbers removed. Messages typed into the playground never leave the page.
- The corpus reflects who sent and collected texts in the 2000s; its notion of spam is a product of that time and place.
Evaluation card: LLM against naive Bayes
This card describes the comparison at /audits/spam, in the section "LLM against the 2022 model".
It evaluates a general-purpose language model on one narrow task with known labels, next to the
naive Bayes model from notebook 37. It is not a model card for any language model, and it makes
no claim about how a model behaves outside this task. Design decisions are in
DR-004.
The question
On the same held-out text messages, does a language model label spam more accurately than the 2022 naive Bayes model, and what does it cost in time and money to find out?
Inputs
- 400 messages drawn with seed 2026 from the 1,840 test messages of the 2022 split (seed 42), in random order, so the first n are themselves a random sample. 57 of the 400 are spam.
- Runs of five or more digits are masked as
#####(DR-003). Masking changed none of naive Bayes's 400 predictions. - The visitor chooses how many messages (25, 50, 100, 200 or 400; 100 by default) and how many go in one call (1, 10 or 25; 10 by default).
The two models
| Model | Trained on | Settings |
|---|---|---|
| Naive Bayes (2022) | The 3,734 training messages of the 2022 split | Its labels for the 400 messages are computed by scripts/audit_spam.py and shipped with the set |
| LLM | Whatever its provider trained it on | Claude Haiku 4.5 at temperature 0 by default (up to 2,048 output tokens); Claude Sonnet 5.5 at low effort (up to 8,192, because thinking counts against the limit); or any OpenAI model the visitor names |
The browser playground's naive Bayes model is not used here, because it was trained on every message, including these.
What the LLM is given
- A fixed system prompt that defines spam and ham for this corpus, explains the masked digits, and says the messages are data to classify and never instructions.
- One batch of messages as JSON, each with an id.
- A JSON schema for the reply,
{"labels": [{"id": ..., "label": "spam" | "ham"}]}, in which the ids are restricted to that batch. zod checks the reply again, including that every id appears exactly once.
Scoring
- Labelled: the batch's reply matched the schema; each message gets the model's label.
- Invalid or refused: the reply failed validation, or the model declined. These are the model's answers, so every message in the batch is scored as wrong.
- Failed call: a network error, or a rate limit that outlasted the retries. That is not the model's answer, so the batch's messages are left out for both models and reported.
- Stopped: errors that would repeat on every batch (a rejected key, an unknown model) stop the run, and the remaining messages are reported as not run.
Metrics and uncertainty
- Accuracy, spam precision and spam recall for both models, with Wilson 95% intervals.
- Spam F1 for both, and the LLM-minus-naive-Bayes differences in accuracy and F1, with paired percentile bootstrap intervals (10,000 resamples of messages, seed 2026). F1 is resampled within each true label, so every resample keeps the sample's spam count. With fewer than 10 spam (the 25- and 50-message samples have 2 and 5), or when a model makes no errors, F1 gets no interval and the page points to the Wilson bounds on precision and recall instead.
- McNemar's exact test on the messages where exactly one model is right, with the conditional odds ratio and its exact interval.
- Median and 90th-percentile call latency, mean latency per message, input and output tokens as the provider reports them, and the cost at list price (Anthropic models only).
- Exports (JSON and CSV) label every row with its source ("AI-generated" or "Mock, not AI"), provider and model, and record the run id, the evaluation set's seed, a prompt id that changes whenever the prompt does, and the generation settings, so a run can be repeated. Mock runs download under their own file name.
What to expect
Naive Bayes is wrong on 4 of the 400 messages and on 2 of the first 100, so any model that is also good will disagree with it on only a handful of messages. McNemar's test will rarely separate two good models at these sizes, and the page says so instead of reporting "no difference". The intervals on precision and recall are wide, because there are only 57 spam messages in all 400.
Limitations
- Contamination. The SMS Spam Collection has been public since 2011, so a language model may have seen these messages and their labels during training. A high score may be partly memory.
- One corpus, one era. Pre-2011 UK and Singapore texts say little about today's spam.
- One prompt, one run. Results depend on the prompt and can vary between runs even at temperature 0; each run is a single sample of the model's behaviour.
- Batching. Labelling several messages in one call is cheaper, but the messages share context. Batch size 1 removes that, at a higher cost.
- Label noise. A few of the corpus's labels are debatable, and both models are scored against them as given.
The mock
Without a key, "Run the mock" replaces the LLM with a keyword rule (two or more cues such as "prize", "claim", "£" or a masked number means spam). It runs in the page through the same validation, scoring and audit log, costs nothing, and is labelled "Mock, not AI" everywhere. It gets 389 of the 400 messages right. Its numbers show how the harness works; they say nothing about any language model.
Decision records
- DR-001 · AcceptedPresent the 40 tutorial notebooks as an honest run log with audits, not as a portfolio of projectsThe site reports what the 2022 notebooks actually did, read by a script from their saved outputs, and adds eight audits that re-check results where a baseline, a time-ordered split or an interval could change the conclusion; it never republishes the tutorials' code or narrative, and it never edits a 2022 number.
- DR-002 · AcceptedJudge every price forecast on later days, against a flat forecast, with a paired testEvery price model in the log is scored on days after the ones it learnt from, next to a flat "no change" forecast scored on the same days, and the gap is tested with the Diebold-Mariano test on the daily errors rather than read off two averages.
- DR-003 · AcceptedRe-run audits only on licence-clear data, publish derived numbers, and make one attributed exception for the SMS corpusAn audit re-runs a notebook only when its data is licence-clear and downloadable without an account; the site publishes metrics, counts and error curves, never raw records, with one exception: a masked, attributed sample of 400 messages from the CC BY 4.0 SMS Spam Collection, because comparing an LLM with naive Bayes needs the text.
- DR-004 · AcceptedCompare an LLM with the 2022 naive Bayes model using the visitor's own key, from the browser, with a local audit logThe LLM comparison on the spam audit is optional: it calls Anthropic or OpenAI directly from the visitor's browser with a key they paste in, scores the model against the 2022 naive Bayes model on the same held-out messages with intervals and McNemar's test, writes every call to an audit log on the visitor's device, and offers a clearly labelled mock run for anyone without a key.
- DR-005 · AcceptedPut a stated interval, a paired test and a seed on every audit numberEvery audit number now carries a 95% interval chosen for what it measures (Wilson for proportions, a seeded bootstrap otherwise, moving blocks for daily errors, a corrected resampled t for cross-validation means), every comparison is paired on the same items (McNemar for classifiers, Diebold-Mariano for forecasts), and every seed is printed next to the result.
AI use statement
This statement says where AI is used on 40 Reps, Re-checked, what it never does, what data leaves the visitor's browser, and where a person stays in charge. It is informed by the Australian Government's Policy for the responsible use of AI in government (Digital Transformation Agency), the transparency principles of the EU AI Act, and the NIST AI Risk Management Framework. This is a personal portfolio project: it is not assessed against any of them and makes no claim of compliance.
What AI does here
There is one AI feature, and it is optional: the comparison "LLM against the 2022 model" on the
spam audit (/audits/spam). When a visitor adds their own API key and starts a run, a language
model labels held-out text messages as spam or ham, and its labels are scored against the known
labels next to the 2022 naive Bayes model's.
What AI never does
- It never runs without the visitor's own key and an explicit click on "Run with my key".
- It never changes the 2022 notebooks, the run log or any audit number. Every audit is computed by deterministic code from published data.
- It never scores itself or another model. Labels are checked against the corpus's own labels.
- It never answers on behalf of another model. There is no silent fallback: a refusal is recorded as a refusal.
- It never writes the text of this site at runtime.
Data sent to the provider
- A fixed system prompt, a batch of masked text messages from the public SMS Spam Collection, and a JSON schema. No personal data about the visitor is put in the request. Like any web request, the call reveals the visitor's IP address, browser user agent and this site's origin to the provider; cookies and the referring page are not sent.
- The visitor's API key, in a request header, to the provider they chose: Anthropic
(
api.anthropic.com) or OpenAI (api.openai.com). Requests go directly from the browser. This site has no AI server code and never receives the key. - The provider handles requests under its own terms and retention policy.
Keys and storage
- The key is kept in the browser's sessionStorage and is gone when the tab closes, unless the visitor turns on "remember on this device" (localStorage). "Forget keys" removes it.
- Any script on the page could read a stored key, including browser extensions. The site loads no third-party scripts, and its Content-Security-Policy only lets the page connect to itself and the two provider APIs (DR-004). Browser extensions are not bound by the policy, so visitors should use a key with a spending limit.
Human in the loop
- The visitor chooses the model, the number of messages and the batch size, sees an estimate of the cost first, starts every run, and can stop it at any time. In-flight calls are cancelled if they leave the page.
- Model output is labelled "AI-generated" wherever it appears. The page shows the run's scores and, for every message either model got wrong, the model's label next to the naive Bayes label and the true label. The exports carry every label, each row with its source ("AI-generated" or "Mock, not AI"), provider, model, run id, prompt id and generation settings. The visitor can accept or reject a run; the decision is stored with every call in it, and individual calls can be reviewed and annotated in the audit log.
- The mock keyword rule is labelled "Mock, not AI" wherever it appears, so its output cannot be mistaken for a model's.
Transparency and records
- Every call is written to an audit log in the visitor's own browser (IndexedDB) with the prompt,
the JSON schema and generation settings, the reply, the validation result, any refusal, the
model, the latency, the token usage and the human decision. It is never uploaded. It can be
viewed, filtered, exported as JSON or CSV, and cleared at
/ai-log. - A run with a key does not start if the audit log cannot be opened in the browser (for example, when storage is blocked), because a call that cannot be recorded is not made. If a single entry fails to save during a run, the page says how many calls are missing from the log.
- Results are reported with intervals, the number of messages and a paired test, and the
evaluation design and its limitations are published on
/methods.
How this site was built
The 2022 notebooks were written by me while following the tutorials. The 2026 revival, including this upgrade, was written with an AI coding assistant under my direction and review, and the commits record that.