Skip to content
40 Reps
36 · Car priceConclusion flips

The perfect score that compared predictions with themselves

Where did the car price model's R² of 1.0 come from?

Short answer

The 2022 call scored the predictions against themselves, which returns 1.0 for any model. Scored against the real prices, the same tree reaches R² 0.85, with a typical miss of $2,372.

What went wrong

The notebook trained a decision tree on 14 numeric columns, predicted the test cars, and then called model.score(X_test, predictions). The second argument should have been the real test prices. Passing the predictions instead asks the model how well it agrees with itself, so the R² is 1.0 every time, for any model.

The script reproduces the call on the UCI data and gets 1.00 on every one of 50 random trees. Then it scores the same trees against the real prices.

The split recipe is the 2022 one (80/20 with seed 42), but it runs on the 195 complete UCI rows rather than the 205 rows of the cleaned copy, so the 39 test cars are not exactly the 2022 ones.

The honest score

Reported in 2022
1.00
predictions against themselves
Same recipe, real prices
0.85
mean R² over 50 trees
Typical miss
$2,372
median price $10,245
Test cars
39
trained on 156
R² on held-out cars (higher is better)
2022 reportself-comparison, meaningless
1.00
Tree, 2022 recipe80/20 split, seed 42
0.85
Tree, 5-fold × 20
0.88
Linear, 5-fold × 20
0.80
Predict the average price
-0.05

The tree's R² on the one split ranges from 0.83 to 0.86 depending on its random seed. Cross-validation 5th to 95th percentile: 0.77 to 0.95.

Intervals and a paired test

Car price model with 95% intervals
MeasureEstimate95% intervalMethodBased on
R², 2022 split (mean of 50 trees)Estimate 0.84895% 0.765 to 0.926Percentile bootstrap (2,000)39 test cars
Typical miss (MAE)Estimate $2,37295% $1,375 to $3,545Percentile bootstrap39 test cars
MAE saved against 'average price' (paired)Estimate $5,21495% $3,381 to $7,203Paired percentile bootstrap39 test cars
Tree R², 5-fold × 20Estimate 0.88295% 0.821 to 0.942Corrected resampled t100 folds

Corrected resampled t

Tree against linear regression, on the same 100 folds

Folds
100
Mean difference
+0.085
95% interval
+0.001 to +0.169
Corrected t
p = 0.05

The tree's advantage of 0.085 R² is real but loosely measured: the interval only just clears zero. On 195 cars, a different random seed for the folds could tell a slightly different story.

Split seed 42 (the 2022 notebook's), tree seeds 0 to 49, cross-validation seed 0, bootstrap seed 2026. The R² interval uses 2,000 resamples because each one rescores all 50 trees. How each interval and test works is set out on the methods page.

What I take from it

The model is decent. An R² near 0.88 means the 14 measurements explain most of the price differences between cars. The problem was only ever the reporting, and a perfect score should have made me suspicious straight away.

Today I would report the error in dollars as well. A typical miss of about $2,372 on a median price of $10,245 is easier to judge than any R².