Four classifiers scored on their own training rows
What do the 2022 scores mean when 360 of the 400 scored rows were used for training?
Short answer
The tree's 0.98 implies 33 of 40 right on the unseen rows, about 82.5%. Logistic regression and the passive-aggressive model both scored exactly 257 of 400, which is what predicting 'no purchase' for everyone would score.
What the notebook did
The notebook split 400 rows into 360 for training and 40 for testing, trained four classifiers on age and estimated salary, then scored each one with score(X, y) on all 400 rows. Nine in ten of the scored rows had already been seen in training.
It also created a support vector machine and a Bernoulli naive Bayes model and never trained or scored them.
Reading the four scores
| Model | 2022 score (all 400 rows) | What it tells us |
|---|---|---|
| Decision tree | 2022 score: 0.9800 | Mostly memory. A fully grown tree gets almost every training row right, so the score says little about new rows. |
| KNN (k = 5) | 2022 score: 0.8750 | Inflated by training rows, but less than the tree, because each prediction averages five neighbours. |
| Logistic regression | 2022 score: 0.6425 | Exactly 257 of 400. Salary in the tens of thousands next to age in years is a classic way to stall the solver, and re-run in a 2022-era scikit-learn it predicts 'no purchase' for every row. |
| Passive aggressive | 2022 score: 0.6425 | The same 257 of 400, for the same reason. |
A score on all rows is a weighted average: 0.9 × training accuracy + 0.1 × test accuracy. A fully grown tree gets every training row right unless two of them share the same age and salary with different answers. The 360 training rows hold one such pair, so the tree tops out at 359 of 360, which the local re-run confirms for every seed. Its 0.98 is 392 of 400 rows, which leaves 33 of the 40 test rows right, about 82.5%. That is still useful, but it is a long way from 98%.
The 0.6425 is exactly 257 out of 400, and 257 people in the table did not buy. Re-run, the logistic regression stops after 3 iterations and predicts “no purchase” for all 400 rows; the passive-aggressive model lands on the same score for most random seeds.
Unmixing the decision tree's 0.98
The score blends 360 training rows with 40 test rows. Choose how many training rows the tree got right and the test accuracy follows. The slider only covers the 8 counts that are possible: below 352 the test rows would need more than 100%, and above 359 the training rows cannot be learnt.
Implied accuracy on the 40 unseen rows: 82.5% (33 of 40)
A decision tree grown with default settings fits its training rows perfectly unless two rows share the same inputs with different labels. The 360 training rows here hold one such pair, so 359 of 360 is the most any tree can reach, and where the slider starts.
How much 40 test rows can say
| Measure | Estimate | 95% interval | Method | Based on |
|---|---|---|---|---|
| Decision tree, implied test accuracy | Estimate 82.5% | 95% 68.1% to 91.3% | Wilson | 33 of 40 |
| Logistic regression, all 400 rows | Estimate 64.3% | 95% 59.4% to 68.8% | Wilson | 257 of 400 |
Even taken at face value, 33 of 40 is consistent with a true accuracy anywhere from 68% to 91%. No paired test is shown: the notebook printed four scores, not predictions, so there is nothing to pair. The logistic regression interval is shown only to make the point that an interval cannot rescue a score measured mostly on training rows.
The 2022 split used random_state 42, so the 40 test rows are known. The tree's own seed was not set; the local re-run tried 200 seeds, every one got 359 of 360 training rows right, and 87.5% of them reproduce the printed 0.98. Nothing here is resampled. How each interval and test works is set out on the methods page.
What I would do now
Score each model on the 40 test rows only, or better, use stratified 10-fold cross-validation so every row is tested once. Standardise age and salary before logistic regression, KNN and the SVM. Print the share of the larger class first, so every score has something to be compared with.