Skip to content
40 Reps
02 · Online payment fraudCan't tell from 2022

99.97% accuracy on a 1-in-775 problem

Does 99.97% accuracy say the fraud model works?

Short answer

99.97% is 381 mistakes on 1,272,524 test rows. That is consistent with catching every fraud, or with missing almost a quarter of them, and the features include balances the dataset's author says must not be used.

What 99.97% is made of

Transactions
6,362,620
Test rows (20%)
1,272,524
Reported accuracy
99.970%
Mistakes
381
exactly, from the accuracy

The notebook printed one number. Working backwards, 99.97006% of 1,272,524 rows is exactly 381 wrong answers. Fraud is about 1 in 775 rows, so the test set holds roughly 1,643 frauds, and a model that never flags anything would already score about 99.87%.

The tree does beat that, with 381 mistakes instead of about 1,643. What accuracy cannot tell me is how those mistakes split between frauds it missed and honest payments it blocked.

Better than never flagging anything?

Fraud model accuracy with 95% intervals
MeasureEstimate95% intervalMethodBased on
Decision tree, as printedEstimate 99.970%95% 99.967% to 99.973%Wilson1,272,143 of 1,272,524
'Never fraud', estimatedEstimate 99.871%95% 99.864% to 99.877%Wilson1,270,881 of 1,272,524

McNemar's test

The tree against 'never fraud', in the split least favourable to the tree

Only tree right
1,643
Only baseline right
381
Exact McNemar
p < 0.0001
Odds ratio, 95%
4.3 (3.9 to 4.8)

The notebook printed no predictions, so the disagreements are not known. But the tree's 381 mistakes can only split between missed frauds and false alarms, and the least favourable split for the tree is all false alarms. Even then it beats the baseline, so on accuracy the tree is better than doing nothing. Whether it is useful is a recall question that accuracy cannot answer, as the explorer below shows.

The baseline assumes about 1,643 frauds in the test rows, from PaySim's reported class count. The 2022 split seed is unknown beyond what the notebook printed, and nothing here is resampled. How each interval and test works is set out on the methods page.

What 99.97% does and does not say

The mistakes are fixed by the accuracy. Whether they are missed frauds or false alarms is not. Move the fraud rate to see how the picture changes.

Share of test transactions that are fraud0.129%

Default 0.129%, the fraud rate of the full PaySim file.

Mistakes the model made
381
Frauds in the test set
1,643
Accuracy of 'never fraud'
99.87%

Fraud recall consistent with 99.97% accuracy: 76.8% to 100.0%

In the worst case all 381 mistakes are missed frauds, so 23.2% of frauds slip through. In the best case they are all false alarms and precision drops to 81.2%.

The features it should not have used

The model used the transaction type, the amount, and the sender's balance before and after the payment. The dataset's own description warns that fraudulent transactions in the simulation were cancelled, so the balance columns should not be used for detection. In other words, the balances partly record the outcome the model is meant to predict.

So even a high recall would be suspect. A fair test drops the balance columns, keeps the class balance in both splits, and reports recall, precision and the precision-recall curve for the fraud class.