Thirty flowers cannot tell 93% from 100%
What did the Iris notebooks' never-printed test score look like, and how sure can it be?
Short answer
The 1-nearest-neighbour model got 30 of 30 test flowers right, but the 95% interval runs from 88.4% to 100%. Over 1,000 other splits only 23% were perfect.
The score nobody printed
- Test flowers
- 30
- 20% of 150
- Correct
- 30 of 30
- 100.0%
- 95% interval (exact)
- 88.4%+
- Clopper-Pearson, up to 100%
- Repeated 5-fold CV
- 95.7%
- sd 3.4%
Notebooks 23 and 15 trained a 1-nearest-neighbour classifier and predicted one flower (Iris-setosa). Neither printed a test score. Running the same split now, the model gets every one of the 30 test flowers right.
A perfect score on 30 flowers sounds conclusive, but it is consistent with a true accuracy anywhere from 88.4% to 100% at 95% confidence. Thirty is simply not many chances to make a mistake.
Other splits, other scores
- 30 of 30100.0%
- 23.1%
- 29 of 3096.7%
- 38.9%
- 28 of 3093.3%
- 27.8%
- 27 of 3090.0%
- 8.9%
- 26 of 3086.7%
- 1.1%
- 25 of 3083.3%
- 0.2%
Only 23% of splits are perfect. The 2022 seed happened to be one of them.
Across 1,000 splits the same model averages 95.8%, close to the cross-validated 95.7%. Using more neighbours moves the cross-validated accuracy by little more than a point, as the table below shows. The honest summary is “about 95%, give or take a few points”.
More neighbours
| Neighbours (k) | 5-fold × 50 accuracy | 95% interval |
|---|---|---|
| 1(2022) | 95.7% | 92.3% to 99.1% |
| 3 | 96.1% | 93.1% to 99.1% |
| 5 | 96.4% | 93.3% to 99.5% |
| 7 | 96.6% | 93.4% to 99.7% |
| 9 | 96.7% | 93.6% to 99.9% |
| 15 | 97.0% | 94.0% to 100.0% |
Each interval is a corrected resampled t over the same 250 folds, which allows for the folds sharing most of their training flowers. The intervals overlap heavily; the paired comparison below asks the sharper question.
Intervals and paired tests
| Measure | Estimate | 95% interval | Method | Based on |
|---|---|---|---|---|
| 1-nearest neighbour, 2022 split | Estimate 100.0% | 95% 88.6% to 100.0% | Wilson | 30 of 30 |
| Same, exact interval | Estimate 100.0% | 95% 88.4% to 100.0% | Clopper-Pearson | 30 of 30 |
| Majority class (Iris-virginica) | Estimate 20.0% | 95% 9.5% to 37.3% | Wilson | 6 of 30 |
| 1-nearest neighbour, 5-fold × 50 | Estimate 95.7% | 95% 92.3% to 99.1% | Corrected resampled t | 250 folds |
McNemar's test
1-nearest neighbour against the majority class, same 30 flowers
- Only model right
- 24
- Only baseline right
- 0
- Exact McNemar
- p < 0.0001
- Odds ratio, 95%
- ∞ (6.0 to ∞)
The model is right on every flower the baseline gets wrong, and never the other way round (p < 0.0001). It learnt something; the open question is only how much.
Corrected resampled t
15 neighbours against 1, on the same 250 folds
- Folds
- 250
- Mean difference
- +0.013
- 95% interval
- −0.020 to +0.046
- Corrected t
- p = 0.44
The table above suggests more neighbours help a little. Paired fold by fold, the gain is 1.3 points with an interval that includes zero, so this data cannot tell the two apart.
Split seed 0 (the 2022 notebooks'), split lottery seeds 0 to 999, cross-validation seed 0. How each interval and test works is set out on the methods page.
Try it with your own numbers
Interval calculator
Change the test size and the number correct. The bars show the range of true accuracies that are consistent with the result at 95% confidence.
Observed accuracy 100.0%. The true accuracy could be anywhere from 88.4% to 100.0% (Clopper-Pearson).