A model that scores each card transaction for review, built on real payment data with a time-ordered holdout. It ranks transactions by risk instead of flagging by rule, so a fixed review budget catches more fraud per hour reviewed.
Measured on 118,108 transactions the model never trained on, held out by time, not by random split. Reviewing the riskiest 1% of transactions catches 25.8% of the fraud in them at 88.7% precision; reviewing the riskiest 2% catches 40.9% at 70.4% precision. From metrics.json in the repo below.
Transactions grouped into 10 equal-sized bins by predicted score. Mean predicted and observed fraud rate line up bin by bin.
| Bin | Transactions | Mean Predicted | Observed Fraud Rate |
|---|---|---|---|
| 1 | 11,810 | 0.027% | 0.178% |
| 2 | 11,811 | 0.124% | 0.195% |
| 3 | 11,811 | 0.178% | 0.390% |
| 4 | 11,811 | 0.224% | 0.644% |
| 5 | 11,811 | 0.274% | 0.771% |
| 6 | 11,810 | 0.662% | 1.109% |
| 7 | 11,811 | 1.075% | 1.363% |
| 8 | 11,811 | 1.880% | 2.151% |
| 9 | 11,811 | 4.260% | 3.920% |
| 10 | 11,811 | 25.634% | 23.690% |
Pick a real holdout transaction and score it against the live model.
Examples are real holdout transactions the model never trained on. Scores run against the live model API.
The model, the training pipeline, and the scoring service: github.com/alphan-ml/fraud-radar
A fixed dollar threshold treats a $30 charge and a $3,000 charge the same. The model scores each transaction on the pattern of the whole card, not one field.
Calibration ties the number to reality: a transaction scored at 20% is a fraud roughly 1 time in 5, checked bin by bin against what actually happened.
Review 1% of transactions and catch a quarter of the fraud at high precision, or review 2% and catch closer to half. The team sets the dial; the model supplies the ranking.