Water quality: the model the notebook never trained
Can these measurements tell safe drinking water from unsafe, and how often is 'safe' wrong?
Short answer
'Always unsafe' already scores 59.7%. The best model reaches 68.8%, and about 32% of the samples it calls safe are not.
Where the notebook stopped
- Samples
- 3,276
- Dropped by dropna
- 1,265
- 38.6% of the data
- Left: safe / unsafe
- 811 / 1,200
- 'Always unsafe' scores
- 59.7%
The 2022 notebook dropped every row with a gap (491 missing pH values, 781 missing sulfate and 162 missing trihalomethanes), drew the class balance twice and ended. No model was trained, so the question “can these measurements tell safe water from unsafe?” was never answered.
The first thing to know is the floor. On the rows that survive, calling every sample unsafe is right 59.7% of the time. Any model has to beat that to be worth anything.
Two models against the floor
- Always 'unsafe'
- 59.7%
- Logistic regression
- 59.8%
- Random forest
- 68.8%
- Random forest, all rowsmedian-imputed, where the floor is 61.0%
- 67.9%
Logistic regression is no better than the floor. Its ROC AUC is 0.48, which is a coin toss, and it calls only 0.7% of the safe samples safe. The relationship between these measurements and potability is not a straight line.
The random forest does find something: 68.8% accuracy and an AUC of 0.71. But at the default cut-off, 32% of the samples it calls safe are unsafe, and it misses 58% of the safe ones. Imputing the gaps instead of dropping them keeps all 3,276 samples and gives a similar result, so dropna was not what held the model back.
How sure, and better than the floor?
| Measure | Estimate | 95% interval | Method | Based on |
|---|---|---|---|---|
| 'Always unsafe' | Estimate 59.7% | 95% 57.5% to 61.8% | Wilson | 1,200 of 2,011 |
| Logistic regression: accuracy | Estimate 59.8% | 95% 57.7% to 61.9% | Wilson | 1,203 of 2,011 |
| Logistic regression: ROC AUC | Estimate 0.483 | 95% 0.455 to 0.509 | Percentile bootstrap (2,000) | 2,011 samples |
| Random forest: accuracy | Estimate 68.8% | 95% 66.8% to 70.8% | Wilson | 1,384 of 2,011 |
| Random forest: 'safe' calls that are safe | Estimate 68.3% | 95% 64.1% to 72.2% | Wilson | 344 of 504 |
| Random forest: safe samples called safe | Estimate 42.4% | 95% 39.1% to 45.8% | Wilson | 344 of 811 |
| Random forest: ROC AUC | Estimate 0.707 | 95% 0.685 to 0.730 | Percentile bootstrap (2,000) | 2,011 samples |
| Forest on all rows (imputed): accuracy | Estimate 67.9% | 95% 66.3% to 69.5% | Wilson | 2,224 of 3,276 |
McNemar's test
Logistic regression against 'always unsafe'
- Only logistic right
- 6
- Only baseline right
- 3
- Exact McNemar
- p = 0.51
- Odds ratio, 95%
- 2.0 (0.4 to 12.4)
The two disagree on only 9 of 2,011 samples, and the split between them is what chance would give (p = 0.51). Logistic regression is the floor with extra steps.
McNemar's test
Random forest against logistic regression
- Only forest right
- 340
- Only logistic right
- 159
- Exact McNemar
- p < 0.0001
- Odds ratio, 95%
- 2.1 (1.8 to 2.6)
On the same out-of-fold predictions the forest is right on 340 samples the linear model gets wrong, against 159 the other way (p < 0.0001). Against ‘always unsafe’ directly, McNemar's exact test gives p < 0.0001 for the forest.
Fold seed 0, forest seed 0, bootstrap seed 2026. The ROC AUC intervals are bootstrapped by the script with 2,000 resamples, because each resample recomputes the AUC over thousands of samples. How each interval and test works is set out on the methods page.
Choosing when to say 'safe'
For drinking water the costly mistake is calling unsafe water safe. Raising the threshold makes the forest more careful, at the price of calling fewer samples safe at all. Move it and watch the unsafe samples that still get called safe.
Random forest: probability of 'safe', out-of-fold
Each sample is scored by a forest that never saw it during training. Bars show how each true class spreads across the scores.
- Really safe (above the line, share of that class)
- Really unsafe (below the line)
- Precision for safe
- 68.3%
- of the safe calls are right
- Recall for safe
- 42.4%
- of the real safe found
- Accuracy
- 68.8%
- Called safe
- 504
- of 2,011 rows
The dangerous cell is unsafe water called safe. Even at strict thresholds a few remain, which is why a model like this could only ever be a screening aid next to a laboratory test.