Spam detection holds up
The 2022 notebook printed no score at all. Is the spam model any good?
Short answer
Yes. It scores 98.4% against 86.1% for 'always ham', catching 96% of spam with 93% precision. Ten-fold cross-validation agrees.
Scoring the 2022 pipeline
- Accuracy
- 98.4%
- 1,840 test messages
- 'Always ham' scores
- 86.1%
- Spam precision
- 93.1%
- spam calls that are right
- Spam recall
- 95.7%
- spam that gets caught
The notebook built word counts, split two thirds for training, trained multinomial naive Bayes and tested one message I typed. It never printed a score. Re-running the same pipeline with the same one-third split and seed, the model is right 98.4% of the time, well above the 86.1% you get by calling everything ham. It still reproduces the 2022 verdict on my test message: spam.
The re-run uses the UCI copy, which has 2 more messages than the 5,572 the notebook read, so its 1,840 test messages are not exactly the 2022 ones. Same recipe, slightly different rows.
18 genuine messages out of 1,585 would have been sent to the spam folder. That is the number I would watch, because a lost message from a friend costs more than one extra spam.
Does it hold on other splits?
- Always 'ham'
- 86.6%
- Accuracy
- 98.7%
- Spam precision
- 97.2%
- Spam recall
- 92.9%
- Spam F1
- 95.0%
Fold-to-fold spread: accuracy 97.7% to 99.3%, recall 88.0% to 97.3%. Here the vocabulary is learnt inside each fold, unlike 2022.
Yes. Cross-validation gives 98.7% accuracy with spam precision of 97.2% and recall of 92.9%. This is the one audit where the 2022 result survives intact. Naive Bayes on word counts is a strong, honest baseline for short spam.
How sure, and better than what?
| Measure | Estimate | 95% interval | Method | Based on |
|---|---|---|---|---|
| Accuracy, 2022 split | Estimate 98.4% | 95% 97.7% to 98.9% | Wilson | 1,811 of 1,840 |
| Spam precision | Estimate 93.1% | 95% 89.4% to 95.6% | Wilson | 244 of 262 |
| Spam recall | Estimate 95.7% | 95% 92.4% to 97.6% | Wilson | 244 of 255 |
| Spam F1 | Estimate 0.944 | 95% 0.923 to 0.963 | Percentile bootstrap, stratified by true label | 1,840 messages |
| 'Always ham' | Estimate 86.1% | 95% 84.5% to 87.6% | Wilson | 1,585 of 1,840 |
| Accuracy, 10-fold CV | Estimate 98.7% | 95% 98.1% to 99.3% | Corrected resampled t | 10 folds |
| Spam F1, 10-fold CV | Estimate 0.950 | 95% 0.928 to 0.972 | Corrected resampled t | 10 folds |
McNemar's test
Naive Bayes against 'always ham' on the same 1,840 messages
- Only naive Bayes right
- 244
- Only 'always ham' right
- 18
- Exact McNemar
- p < 0.0001
- Odds ratio, 95%
- 13.6 (8.4 to 23.3)
The model alone is right on the 244 spam messages it catches; the baseline alone is right on the 18 genuine messages the model flags. That is +12.3 percentage points of accuracy (95% interval +10.7 to +14.0, paired bootstrap; p < 0.0001). The 2022 verdict holds with room to spare.
Split seed 42 (the 2022 notebook's), cross-validation seed 0, bootstrap seed 2026 with 10,000 resamples. Cross-validation intervals use the corrected resampled t, because the ten folds share most of their training messages. How each interval and test works is set out on the methods page.
Trading missed spam for lost messages
Spam score (log-odds), out-of-fold
Naive Bayes is very confident, so scores are shown as log-odds. Zero is the default cut-off. Raising it sends fewer genuine messages to spam.
- Really spam (above the line, share of that class)
- Really ham (below the line)
- Bar heights use a square-root scale so thin bins stay visible.
- Precision for spam
- 97.2%
- of the spam calls are right
- Recall for spam
- 92.9%
- of the real spam found
- Accuracy
- 98.7%
- Called spam
- 714
- of 5,574 rows
False alarms are genuine messages filtered as spam. A few points of recall are usually worth giving up to keep them near zero.
Try the model
Try the spam model
The same multinomial naive Bayes model, ported to TypeScript and running in your browser. Nothing you type is sent anywhere.
Loading the model (about 100 KB)…
The browser model is the same naive Bayes trained on all 5,574 messages, with one change: tokens containing five or more digits in a row (phone numbers and short codes) are left out of the published vocabulary. A parity test checks the TypeScript port against scikit-learn's own probabilities.
LLM against the 2022 model
Naive Bayes on word counts is a strong baseline for short spam. Would a general-purpose language model do better on the same held-out messages, and at what cost? This optional comparison uses your own API key, calls the provider straight from your browser, and scores both models on identical messages with intervals and a paired test. Without a key you can run a mock, a keyword rule clearly labelled as such, to see how the harness works.
The 400 messages (57 spam) are a random sample of the 2022 test split (seed 2026) from the SMS Spam Collection by Almeida, Gómez Hidalgo and Yamakami, CC BY 4.0, with runs of five or more digits masked as #####. Naive Bayes gets 396 of them right. Every call is recorded in your AI audit log; the design and its limits are on the methods page.
Loading the evaluation set…