WEBVTT

NOTE Spam: classical vs LLM: captions match the on-screen steps.

1
00:00:00.200 --> 00:00:07.337
Step 1 of 8. SMS spam: the 2022 naive Bayes model holds up at 98.4% accuracy (95% CI 97.7% to 98.9%).

2
00:00:07.337 --> 00:00:11.522
Step 2 of 8. Against 'always ham' on the same 1,840 messages: McNemar's exact test, p < 0.0001.

3
00:00:11.522 --> 00:00:18.970
Step 3 of 8. LLM against naive Bayes: the same held-out messages, your own API key, called from your browser.

4
00:00:18.970 --> 00:00:32.856
Step 4 of 8. AI settings: Anthropic or OpenAI, a cheap model by default, the key kept in this browser only.

5
00:00:32.856 --> 00:00:41.335
Step 5 of 8. No key? Run the mock: a keyword rule labelled 'Mock, not AI' (mocked run for illustration).

6
00:00:41.335 --> 00:00:47.454
Step 6 of 8. Same 100 messages: 98.0% each, with Wilson intervals and a paired bootstrap difference.

7
00:00:47.454 --> 00:00:53.388
Step 7 of 8. McNemar's test: 4 disagreements, p = 1.00, too few to tell the two apart.

8
00:00:53.388 --> 00:01:08.000
Step 8 of 8. Accept or reject the run: every call lands in the AI audit log, exportable as JSON or CSV.
