The perfect score that compared predictions with themselves
Where did the car price model's R² of 1.0 come from?
Short answer
The 2022 call scored the predictions against themselves, which returns 1.0 for any model. Scored against the real prices, the same tree reaches R² 0.85, with a typical miss of $2,372.
What went wrong
The notebook trained a decision tree on 14 numeric columns, predicted the test cars, and then called model.score(X_test, predictions). The second argument should have been the real test prices. Passing the predictions instead asks the model how well it agrees with itself, so the R² is 1.0 every time, for any model.
The script reproduces the call on the UCI data and gets 1.00 on every one of 50 random trees. Then it scores the same trees against the real prices.
The split recipe is the 2022 one (80/20 with seed 42), but it runs on the 195 complete UCI rows rather than the 205 rows of the cleaned copy, so the 39 test cars are not exactly the 2022 ones.
The honest score
- Reported in 2022
- 1.00
- predictions against themselves
- Same recipe, real prices
- 0.85
- mean R² over 50 trees
- Typical miss
- $2,372
- median price $10,245
- Test cars
- 39
- trained on 156
- 2022 reportself-comparison, meaningless
- 1.00
- Tree, 2022 recipe80/20 split, seed 42
- 0.85
- Tree, 5-fold × 20
- 0.88
- Linear, 5-fold × 20
- 0.80
- Predict the average price
- -0.05
The tree's R² on the one split ranges from 0.83 to 0.86 depending on its random seed. Cross-validation 5th to 95th percentile: 0.77 to 0.95.
Intervals and a paired test
| Measure | Estimate | 95% interval | Method | Based on |
|---|---|---|---|---|
| R², 2022 split (mean of 50 trees) | Estimate 0.848 | 95% 0.765 to 0.926 | Percentile bootstrap (2,000) | 39 test cars |
| Typical miss (MAE) | Estimate $2,372 | 95% $1,375 to $3,545 | Percentile bootstrap | 39 test cars |
| MAE saved against 'average price' (paired) | Estimate $5,214 | 95% $3,381 to $7,203 | Paired percentile bootstrap | 39 test cars |
| Tree R², 5-fold × 20 | Estimate 0.882 | 95% 0.821 to 0.942 | Corrected resampled t | 100 folds |
Corrected resampled t
Tree against linear regression, on the same 100 folds
- Folds
- 100
- Mean difference
- +0.085
- 95% interval
- +0.001 to +0.169
- Corrected t
- p = 0.05
The tree's advantage of 0.085 R² is real but loosely measured: the interval only just clears zero. On 195 cars, a different random seed for the folds could tell a slightly different story.
Split seed 42 (the 2022 notebook's), tree seeds 0 to 49, cross-validation seed 0, bootstrap seed 2026. The R² interval uses 2,000 resamples because each one rescores all 50 trees. How each interval and test works is set out on the methods page.
What I take from it
The model is decent. An R² near 0.88 means the 14 measurements explain most of the price differences between cars. The problem was only ever the reporting, and a perfect score should have made me suspicious straight away.
Today I would report the error in dollars as well. A typical miss of about $2,372 on a median price of $10,245 is easier to judge than any R².