Decision record · DR-004
Compare an LLM with the 2022 naive Bayes model using the visitor's own key, from the browser, with a local audit log
- Status
- Accepted
- Date
- 2026-10
- Applies to
- /audits/spam#llm, /ai-log, web/src/lib/ai, web/src/lib/security/csp.ts, web/next.config.ts
Decision in one line
The LLM comparison on the spam audit is optional: it calls Anthropic or OpenAI directly from the visitor's browser with a key they paste in, scores the model against the 2022 naive Bayes model on the same held-out messages with intervals and McNemar's test, writes every call to an audit log on the visitor's device, and offers a clearly labelled mock run for anyone without a key.
Context
The spam audit is the one place in the log where the 2022 model holds up: naive Bayes on word counts scores 98.4% on the held-out messages. That makes it a fair test of a question people now ask about every classical model: would a general-purpose language model do better, and at what cost? The site is static and has no budget for AI calls. A shared key held by the site would cost money with every visit and would need abuse protection and secret management.
Decision
- The visitor picks Anthropic (default: Claude Haiku 4.5 at temperature 0; Claude Sonnet 5.5, which always thinks, at low effort, as the option) or OpenAI (model id is free text) and pastes their own key. The key is kept in sessionStorage, moves to localStorage only if they turn on "remember on this device", and "Forget keys" clears both.
- Requests go from the browser straight to the provider, with Anthropic's
anthropic-dangerous-direct-browser-accessheader. The site has no AI server code and never receives the key. Every page is served with a Content-Security-Policy whoseconnect-srclists only this site and the two provider APIs, so a script injected into the page could not send the key anywhere else withfetch. - The model sees a fixed system prompt that defines spam and ham, says the messages are data and never instructions, and a batch of messages as JSON. Its reply is held to a JSON schema whose ids are exactly that batch's, and zod checks it again, including that every id appears once.
- Scoring rules were fixed before any run. A reply that fails validation, or a refusal, is the model's answer and counts as wrong for every message in its batch. A failed call (network, or a rate limit after retries) is not the model's answer: its messages are left out for both models and reported. Errors that would repeat on every batch (bad key, unknown model) stop the run.
- Both models are scored on the same messages: accuracy, spam precision and recall with Wilson intervals, F1 and the paired differences with a seeded percentile bootstrap (resampled within each true label for F1, and no F1 interval below 10 spam or when a model makes no errors), and McNemar's exact test on the messages where they disagree. Latency, tokens and the list-price cost are reported next to the scores.
- Every call is appended to an audit log in IndexedDB (time, feature, provider, model, full
prompt, schema and settings, raw and parsed output, validation result, stop reason, latency,
tokens, human decision and note), scrubbed for the key before it is stored, and viewable and
exportable as JSON or CSV at
/ai-log. The visitor can accept or reject a run, and the decision is stored against every call in it. A run with a key does not start when the log cannot be opened, and the page reports any call whose entry failed to save. - Without a key, "Run the mock" swaps in a keyword rule that runs in the page through the same pipeline, validation and audit log. It is labelled "Mock, not AI" everywhere it appears, and its log entries say "no AI called".
- Every model output on the site carries an "AI-generated" label, and every exported row says whether it is "AI-generated" or "Mock, not AI", with the provider, model, run id, prompt id and generation settings. Mock runs download under their own file name.
Options considered
- A server route with my key. Every visit would cost me money and need rate limits.
- A server route with the visitor's key. The key would pass through a server I run.
- Browser-direct calls with the visitor's key (chosen).
- Precomputed results from one paid run. Cheap to show, but visitors could not reproduce or vary it, and it would not exercise the audit trail.
- No AI feature. Simplest, but it leaves out the evaluation and governance work the upgrade is meant to show.
Why
Option 3 costs nothing to host, keeps the key between the visitor and their provider, and keeps the audit record on the visitor's device, where they can inspect and export it. Scoring against known labels, next to a model whose labels are fixed in advance, means there is no judge model and no rubric to argue about. The mock lets anyone see the harness, the statistics and the audit log work without spending anything, and labelling it plainly keeps it from being mistaken for evidence about a model.
What happened
- The adapters, validation, retries, scoring rules, audit log and export are covered by unit
tests with mocked
fetchresponses. No paid run is recorded in this repository: a run needs a key and spends money, so the page reports only runs made in the visitor's own browser. - The comparison is hard to win and hard to lose. On the first 100 messages of the evaluation set naive Bayes is wrong on 2, so even a perfect model would disagree with it on very few messages, and McNemar's test will rarely separate them at that size. The page says so next to the result rather than leaving a non-significant p-value to be read as "no difference".
- The mock keyword rule gets 389 of the 400 messages right (97.3%) against naive Bayes's 396. A crude rule doing that well is a reminder of how easy much of this corpus is.
- A static Content-Security-Policy still allows inline scripts, because the pages are prerendered without per-request nonces. It limits where a script could send data, not what it could read, and browser extensions are not bound by it, so the advice to use a key with a spending limit stands.
What I'd change
- Run each batch several times to measure run-to-run variation, as well as the mean.
- Try a second prompt (for example with a few labelled examples) and report the difference as a paired comparison on the same messages.
- Add hand-written messages the model cannot have seen, to check for contamination (DR-003).
- Record the provider's request id in each audit entry, so an entry can be matched to the provider's own usage logs.