Skip to content
40 Reps
24 · Drinking water qualityWeaker than it looked

Water quality: the model the notebook never trained

Can these measurements tell safe drinking water from unsafe, and how often is 'safe' wrong?

Short answer

'Always unsafe' already scores 59.7%. The best model reaches 68.8%, and about 32% of the samples it calls safe are not.

Where the notebook stopped

Samples
3,276
Dropped by dropna
1,265
38.6% of the data
Left: safe / unsafe
811 / 1,200
'Always unsafe' scores
59.7%

The 2022 notebook dropped every row with a gap (491 missing pH values, 781 missing sulfate and 162 missing trihalomethanes), drew the class balance twice and ended. No model was trained, so the question “can these measurements tell safe water from unsafe?” was never answered.

The first thing to know is the floor. On the rows that survive, calling every sample unsafe is right 59.7% of the time. Any model has to beat that to be worth anything.

Two models against the floor

Accuracy, 5-fold cross-validation (higher is better)
Always 'unsafe'
59.7%
Logistic regression
59.8%
Random forest
68.8%
Random forest, all rowsmedian-imputed, where the floor is 61.0%
67.9%

Logistic regression is no better than the floor. Its ROC AUC is 0.48, which is a coin toss, and it calls only 0.7% of the safe samples safe. The relationship between these measurements and potability is not a straight line.

The random forest does find something: 68.8% accuracy and an AUC of 0.71. But at the default cut-off, 32% of the samples it calls safe are unsafe, and it misses 58% of the safe ones. Imputing the gaps instead of dropping them keeps all 3,276 samples and gives a similar result, so dropna was not what held the model back.

How sure, and better than the floor?

Water models with 95% intervals
MeasureEstimate95% intervalMethodBased on
'Always unsafe'Estimate 59.7%95% 57.5% to 61.8%Wilson1,200 of 2,011
Logistic regression: accuracyEstimate 59.8%95% 57.7% to 61.9%Wilson1,203 of 2,011
Logistic regression: ROC AUCEstimate 0.48395% 0.455 to 0.509Percentile bootstrap (2,000)2,011 samples
Random forest: accuracyEstimate 68.8%95% 66.8% to 70.8%Wilson1,384 of 2,011
Random forest: 'safe' calls that are safeEstimate 68.3%95% 64.1% to 72.2%Wilson344 of 504
Random forest: safe samples called safeEstimate 42.4%95% 39.1% to 45.8%Wilson344 of 811
Random forest: ROC AUCEstimate 0.70795% 0.685 to 0.730Percentile bootstrap (2,000)2,011 samples
Forest on all rows (imputed): accuracyEstimate 67.9%95% 66.3% to 69.5%Wilson2,224 of 3,276

McNemar's test

Logistic regression against 'always unsafe'

Only logistic right
6
Only baseline right
3
Exact McNemar
p = 0.51
Odds ratio, 95%
2.0 (0.4 to 12.4)

The two disagree on only 9 of 2,011 samples, and the split between them is what chance would give (p = 0.51). Logistic regression is the floor with extra steps.

McNemar's test

Random forest against logistic regression

Only forest right
340
Only logistic right
159
Exact McNemar
p < 0.0001
Odds ratio, 95%
2.1 (1.8 to 2.6)

On the same out-of-fold predictions the forest is right on 340 samples the linear model gets wrong, against 159 the other way (p < 0.0001). Against ‘always unsafe’ directly, McNemar's exact test gives p < 0.0001 for the forest.

Fold seed 0, forest seed 0, bootstrap seed 2026. The ROC AUC intervals are bootstrapped by the script with 2,000 resamples, because each resample recomputes the AUC over thousands of samples. How each interval and test works is set out on the methods page.

Choosing when to say 'safe'

For drinking water the costly mistake is calling unsafe water safe. Raising the threshold makes the forest more careful, at the price of calling fewer samples safe at all. Move it and watch the unsafe samples that still get called safe.

Random forest: probability of 'safe', out-of-fold

Each sample is scored by a forest that never saw it during training. Bars show how each true class spreads across the scores.

  • Really safe (above the line, share of that class)
  • Really unsafe (below the line)
P(safe) ≥ 0.50 → safe0.001.00
Decision thresholdP(safe) ≥ 0.50
Precision for safe
68.3%
of the safe calls are right
Recall for safe
42.4%
of the real safe found
Accuracy
68.8%
Called safe
504
of 2,011 rows
Actual class, by predictionCalled safeCalled unsafe
Really safe344safe, called safe467safe, called unsafe
Really unsafe160UNSAFE, called safe1,040unsafe, called unsafe

The dangerous cell is unsafe water called safe. Even at strict thresholds a few remain, which is why a model like this could only ever be a screening aid next to a laboratory test.