Skip to content
40 Reps

Thirty flowers cannot tell 93% from 100%

What did the Iris notebooks' never-printed test score look like, and how sure can it be?

Short answer

The 1-nearest-neighbour model got 30 of 30 test flowers right, but the 95% interval runs from 88.4% to 100%. Over 1,000 other splits only 23% were perfect.

The score nobody printed

Test flowers
30
20% of 150
Correct
30 of 30
100.0%
95% interval (exact)
88.4%+
Clopper-Pearson, up to 100%
Repeated 5-fold CV
95.7%
sd 3.4%

Notebooks 23 and 15 trained a 1-nearest-neighbour classifier and predicted one flower (Iris-setosa). Neither printed a test score. Running the same split now, the model gets every one of the 30 test flowers right.

A perfect score on 30 flowers sounds conclusive, but it is consistent with a true accuracy anywhere from 88.4% to 100% at 95% confidence. Thirty is simply not many chances to make a mistake.

Other splits, other scores

Correct out of 30 across 1,000 random splits (share of splits)
30 of 30100.0%
23.1%
29 of 3096.7%
38.9%
28 of 3093.3%
27.8%
27 of 3090.0%
8.9%
26 of 3086.7%
1.1%
25 of 3083.3%
0.2%

Only 23% of splits are perfect. The 2022 seed happened to be one of them.

Across 1,000 splits the same model averages 95.8%, close to the cross-validated 95.7%. Using more neighbours moves the cross-validated accuracy by little more than a point, as the table below shows. The honest summary is “about 95%, give or take a few points”.

More neighbours

Cross-validated accuracy by number of neighbours, with 95% corrected resampled t intervals
Neighbours (k)5-fold × 50 accuracy95% interval
1(2022)95.7%92.3% to 99.1%
396.1%93.1% to 99.1%
596.4%93.3% to 99.5%
796.6%93.4% to 99.7%
996.7%93.6% to 99.9%
1597.0%94.0% to 100.0%

Each interval is a corrected resampled t over the same 250 folds, which allows for the folds sharing most of their training flowers. The intervals overlap heavily; the paired comparison below asks the sharper question.

Intervals and paired tests

Iris results with 95% intervals
MeasureEstimate95% intervalMethodBased on
1-nearest neighbour, 2022 splitEstimate 100.0%95% 88.6% to 100.0%Wilson30 of 30
Same, exact intervalEstimate 100.0%95% 88.4% to 100.0%Clopper-Pearson30 of 30
Majority class (Iris-virginica)Estimate 20.0%95% 9.5% to 37.3%Wilson6 of 30
1-nearest neighbour, 5-fold × 50Estimate 95.7%95% 92.3% to 99.1%Corrected resampled t250 folds

McNemar's test

1-nearest neighbour against the majority class, same 30 flowers

Only model right
24
Only baseline right
0
Exact McNemar
p < 0.0001
Odds ratio, 95%
∞ (6.0 to ∞)

The model is right on every flower the baseline gets wrong, and never the other way round (p < 0.0001). It learnt something; the open question is only how much.

Corrected resampled t

15 neighbours against 1, on the same 250 folds

Folds
250
Mean difference
+0.013
95% interval
−0.020 to +0.046
Corrected t
p = 0.44

The table above suggests more neighbours help a little. Paired fold by fold, the gain is 1.3 points with an interval that includes zero, so this data cannot tell the two apart.

Split seed 0 (the 2022 notebooks'), split lottery seeds 0 to 999, cross-validation seed 0. How each interval and test works is set out on the methods page.

Try it with your own numbers

Interval calculator

Change the test size and the number correct. The bars show the range of true accuracies that are consistent with the result at 95% confidence.

Observed accuracy 100.0%. The true accuracy could be anywhere from 88.4% to 100.0% (Clopper-Pearson).

95% intervals
Wilson score88.6% to 100.0%
Clopper-Pearson (exact)88.4% to 100.0%
80%85%90%95%100%