99.97% accuracy on a 1-in-775 problem
Does 99.97% accuracy say the fraud model works?
Short answer
99.97% is 381 mistakes on 1,272,524 test rows. That is consistent with catching every fraud, or with missing almost a quarter of them, and the features include balances the dataset's author says must not be used.
What 99.97% is made of
- Transactions
- 6,362,620
- Test rows (20%)
- 1,272,524
- Reported accuracy
- 99.970%
- Mistakes
- 381
- exactly, from the accuracy
The notebook printed one number. Working backwards, 99.97006% of 1,272,524 rows is exactly 381 wrong answers. Fraud is about 1 in 775 rows, so the test set holds roughly 1,643 frauds, and a model that never flags anything would already score about 99.87%.
The tree does beat that, with 381 mistakes instead of about 1,643. What accuracy cannot tell me is how those mistakes split between frauds it missed and honest payments it blocked.
Better than never flagging anything?
| Measure | Estimate | 95% interval | Method | Based on |
|---|---|---|---|---|
| Decision tree, as printed | Estimate 99.970% | 95% 99.967% to 99.973% | Wilson | 1,272,143 of 1,272,524 |
| 'Never fraud', estimated | Estimate 99.871% | 95% 99.864% to 99.877% | Wilson | 1,270,881 of 1,272,524 |
McNemar's test
The tree against 'never fraud', in the split least favourable to the tree
- Only tree right
- 1,643
- Only baseline right
- 381
- Exact McNemar
- p < 0.0001
- Odds ratio, 95%
- 4.3 (3.9 to 4.8)
The notebook printed no predictions, so the disagreements are not known. But the tree's 381 mistakes can only split between missed frauds and false alarms, and the least favourable split for the tree is all false alarms. Even then it beats the baseline, so on accuracy the tree is better than doing nothing. Whether it is useful is a recall question that accuracy cannot answer, as the explorer below shows.
The baseline assumes about 1,643 frauds in the test rows, from PaySim's reported class count. The 2022 split seed is unknown beyond what the notebook printed, and nothing here is resampled. How each interval and test works is set out on the methods page.
What 99.97% does and does not say
The mistakes are fixed by the accuracy. Whether they are missed frauds or false alarms is not. Move the fraud rate to see how the picture changes.
Default 0.129%, the fraud rate of the full PaySim file.
- Mistakes the model made
- 381
- Frauds in the test set
- 1,643
- Accuracy of 'never fraud'
- 99.87%
Fraud recall consistent with 99.97% accuracy: 76.8% to 100.0%
In the worst case all 381 mistakes are missed frauds, so 23.2% of frauds slip through. In the best case they are all false alarms and precision drops to 81.2%.
The features it should not have used
The model used the transaction type, the amount, and the sender's balance before and after the payment. The dataset's own description warns that fraudulent transactions in the simulation were cancelled, so the balance columns should not be used for detection. In other words, the balances partly record the outcome the model is meant to predict.
So even a high recall would be suspect. A fair test drops the balance columns, keeps the class balance in both splits, and reports recall, precision and the precision-recall curve for the fraud class.