Decision models

Four fraud detectors, open on Hugging Face

We put four fraud detectors on Hugging Face: System One, our decision model, stacked on gradient-boosted trees, one for each Fraud Dataset Benchmark set we can publish. This post walks through each model card, losses included.

STEAV8 min read

What we published

Four repositories under STEAV on Hugging Face, one per benchmark set, gathered in one collection. Each holds everything needed to score its benchmark: the gradient-boosted tree model, the System One adapter and decision head, the inference code, a script that rebuilds the official split with the benchmark's own loader and recomputes the metrics, and a model card. Weights and code are Apache-2.0. The few data files that ship inside a model keep their own terms, and each card lists them file by file.

The Fraud Dataset Benchmark has nine sets. The other five stay internal: their data carries non-commercial or research-only terms, and we do not publish models trained on data like that.

Why there is a tree underneath

In September we wrote that System One is not a tabular model: where the evidence is numbers and codes, gradient boosting on the same rows ranks better, and the way through is to hand the encoder the tree's score as one more field. These models do exactly that. A gradient-boosted tree model scores each row from engineered features. System One then reads that score, its percentile among the training scores, a few readable engineered fields and the original fields as JSON evidence, and answers one yes-or-no question with a calibrated probability. On IP addresses, System One reading the address alone scored 0.740 AUROC in September; stacked on its tree, it scores 0.9481.

On two sets the released output is neither System One alone nor the tree alone, but a logistic blend of the two, fitted on System One's holdout split: on IP addresses because System One alone trailed its tree, and on the ULB card data because the blend scored highest. Every card prints all three, so the choice is visible.

The results, side by side

Every comparison below is on the same test rows. Each chart sets our released model against three things: the best result published with the benchmark in 2022, when its authors ran AutoGluon, H2O, Auto-sklearn and Amazon Fraud Detector; AutoGluon 1.6.3, today's release of the AutoML system that set or shared three of those records, which we re-ran with the benchmark's own script and settings; and, for the ULB card data, the two recent papers that scored the same test window, Sang (2026) and Breskuvienė and Dzemyda (2024). We found no newer published result on these test splits. Amazon Fraud Detector cannot be re-run: it closed to new customers in November 2025, and AWS now points them to AutoGluon.

Figure 1Fraud caught at 1% false alarms, on the same test rows.
Scores on each set's official test rows; whiskers are 95% intervals (class-stratified bootstrap, 1,000 resamples) where we computed them. Best published: the benchmark's paper (AUROC) and its README results table (recall). AutoGluon 1.6.3 ran the benchmark's own script and settings (best_quality, one hour, the same feature columns). Our stored copy of the ULB data keeps each number as exact decimal text, so we converted those 29 columns back to numbers first; left as text, AutoGluon reads every value as a category. On IP addresses the script gives AutoGluon only the raw address string, nearly every one unique, so it scores at chance.
Figure 2AUROC, on the same test rows.
Scores on each set's official test rows; whiskers are 95% intervals (class-stratified bootstrap, 1,000 resamples) where we computed them. Best published: the benchmark's paper (AUROC) and its README results table (recall). AutoGluon 1.6.3 ran the benchmark's own script and settings (best_quality, one hour, the same feature columns). Our stored copy of the ULB data keeps each number as exact decimal text, so we converted those 29 columns back to numbers first; left as text, AutoGluon reads every value as a category. On IP addresses the script gives AutoGluon only the raw address string, nearly every one unique, so it scores at chance.

On fake job postings, IP addresses and simulated card transactions, the released model has the highest AUROC on the chart. It also catches the most fraud at 1% false alarms on the first two. On simulated cards, Amazon Fraud Detector caught all 93 frauds in 2022 to our 91, and that service can no longer be re-run. AutoGluon's 0.50 on IP addresses is not a weak model: the benchmark's script hands it only the raw address, almost every one unique, and nothing derived from it.

The ULB card data is where we lose. H2O's published AUROC of 0.992 and today's AutoGluon at 0.9907 are both ahead of our 0.9867, and AutoGluon catches 67 of the 75 frauds at 1% false alarms to our 65. Our intervals overlap theirs, so with 75 frauds the gap is not settled, but the point estimates are behind and the card says so.

The released models, exactlyEach set's official test split, scored once per model.
Fake job postings (3,576; 187 fake)
AUROC
0.9985
Caught at 1% false alarms
97.3%
95% intervals
AUROC 0.9970–0.9996 · caught 94.7–99.5%
Released output
System One
Best published
AUROC 0.998 · caught 92.5% (AutoGluon)
IP addresses (43,000; 2,997 malicious)
AUROC
0.9481
Caught at 1% false alarms
55.7%
95% intervals
AUROC 0.9441–0.9520 · caught 53.7–57.8%
Released output
Blend of System One and the tree
Best published
AUROC 0.937 · caught 46.6% (AFD OFI)
Simulated card transactions (20,000; 93 fraudulent)
AUROC
0.9994
Caught at 1% false alarms
97.9%
95% intervals
AUROC 0.9989–0.9999 · caught 94.6–100.0%
Released output
System One
Best published
AUROC 0.998 · caught 100.0% (AFD OFI)
ULB card transactions (56,962; 75 fraudulent)
AUROC
0.9867
Caught at 1% false alarms
86.7%
95% intervals
AUROC 0.9748–0.9961 · caught 78.7–93.3%
Released output
Blend of System One and the tree
Best published
AUROC 0.992 (H2O) · caught 88.0% (a three-way tie of AFD OFI, AFD TFI and AutoGluon)
Caught at 1% false alarms is recall at a 1% false-positive rate, computed as the benchmark does. Intervals: class-stratified bootstrap, 1,000 resamples. Published AUROC is from the benchmark's paper, which reports no recall; published recall is from the results table in its README at commit 54cdefa211.

The newest tabular systems have not been run on these test splits, by us or anyone else. On the same data under their own protocols they score as below. Those are different test rows, so we list them for context and do not chart them against ours.

Recent results on the same data, other protocolsContext only. Different test rows, so these are not comparable point for point.
DataRecent AUROCSource and protocol
Fake job postingsLimiX-2 0.9974 · Kumo-Tabular 0.9969 · TabPFN-3.5 0.9956 · tuned XGBoost 0.9891BeyondArena (TabArena), 2026: repeated random three-fold splits over 17,460 postings
ULB card transactionsTabPFN 3.5 0.9902Prior Labs, 2026: a random 80/20 split, which is easier than the benchmark's time-ordered test
Simulated card transactionsXGBoost 0.9963Giusti et al., 2026: the full Sparkov test file of 555,719 rows; the benchmark samples 20,000
IP addressesnone foundThis list appears in no published work outside the benchmark itself.

We picked each released output after the test scores were in, and picking among candidates after the fact can flatter a headline slightly. That is why every card prints every candidate we scored: System One alone, the tree alone and the blend.

Fake job postings

The data is EMSCAD: 17,880 real job ads published through one applicant-tracking platform in 2012 to 2014, 866 of them fraudulent, released by the University of the Aegean for research and for training anti-scam classifiers. The tree blends three families of stackers over 87 features, among them text models and features that ask whether a posting is a near copy of a known fake. System One reads the posting's fields, with long text cut to its first 700 characters, alongside the tree's view, and it is the strongest of the three candidates: 0.9985 AUROC and 97.3% of fake postings caught at 1% false alarms, against AutoGluon's 0.998 and 92.5%.

  • The benchmark rewards memorisation. The split is random, and 78 test postings are exact copies of training postings. On postings with no near copy in training, the released tree recipe scores 0.991 AUROC and 88.3% recall in cross-validation. New scam styles will score lower than the headline.
  • The model carries a recoverable fingerprint of its training data. To compute the near-copy features it ships a hashed vector of every training posting. An independent test rebuilt about 68% of each posting's words and 1,126 exact titles from the repository and a word list. EMSCAD is public already, and the card says plainly what can be rebuilt.

IP blocklist

The data is the CINS Army threat-intelligence list as the benchmark packaged it in 2022: 15,000 listed addresses against 200,000 random IPv4 addresses labelled benign. The tree reads nothing but the address: its octets and prefixes, how crowded its neighbourhood is, how near it sits to known-malicious and known-benign training addresses, and which network owns it. The released blend catches 55.7% of malicious addresses at 1% false alarms with 0.9481 AUROC, against the best published 46.6% and 0.937.

  • The benign class is synthetic. Random addresses are not real traffic. The model largely learns whether an address sits in networks that host listed attackers, so on real traffic its precision will be much lower than the benchmark suggests.
  • Part of the lead comes from later data. The table of network owners is a 2026 snapshot applied to a 2022 list. Refresh it, and retrain, before any real use.
  • The bundled addresses are not a blocklist. They ship so new addresses can be compared with them. A 2022 list is stale, and blocking its addresses today would hit whoever holds them now.

Simulated card transactions

The data is Sparkov, a public generator's simulation of 1,000 customers and 800 merchants, released as public domain. The tree blends LightGBM and CatBoost models over the amount, category, time and customer, plus card profiles built only from each card's transactions older than seven days. System One, the tree and the blend tie on the headline: 0.9994 AUROC and 97.9% caught at 1% false alarms. The published 100% is two more frauds out of 93.

  • Simulated fraud is easy fraud. Sparkov generates a simple pattern: bursts of unusual-category, unusual-amount, mostly night-time spending. These results do not transfer to real card data.
  • A simulator quirk is left out on purpose. Every card has at most one fraud burst. A feature built on that would raise recall here and mean nothing anywhere else, so we did not build it.
  • It needs each card's history. The package bundles the training-period history so the benchmark reproduces; for new transactions you pass each card's history in.

ULB card transactions

The data is two days of European card payments from September 2013, released by the Université Libre de Bruxelles with Worldline: 28 anonymised principal components, the amount and the time, split so the test period follows training. The tree is a bag of 15 XGBoost models, and the released output blends it with System One: 0.9867 AUROC and 86.7% caught at 1% false alarms. That trails the best published 0.992 and 88.0%, though both published values sit inside our intervals, with 75 test frauds making one fraud worth 1.3 points of recall.

  • System One adds nothing here. Its evidence is numbers it cannot interpret, its token window leaves out 11 of the 28 components, and the fitted blend gives it a negative weight, so the output follows the tree. We ship it anyway, for one decision interface across all four sets.
  • Open database, open method. The data is under the Open Database License, so the repository documents how the training data was derived from it, row lists included, as the licence asks.

How we checked them

A model card is only as good as the numbers behind it, so each repository had to pass these checks before it went public:

  • The tree model rebuilds. Refit from scratch, it reproduces its evaluated scores bit for bit on every test row.
  • The evidence is identical. The shipped code reproduces all 123,538 test states that System One reads, byte for byte.
  • System One reproduces. Bit for bit on every test row on the reference GPU, an NVIDIA GB10, with the base model downloaded from the Hub.
  • A clean install works. A fresh environment built from the card's pinned requirements reproduces the tree scores and the evidence exactly.
  • The benchmark rebuilds. The benchmark's own loader, run fresh, recreates the evaluated splits, and eval/reproduce_fdb.py recomputes the metrics end to end on a CPU from that fresh download. It matched the reported AUROC and recall to four decimals on all four sets.
  • Three independent reviews. Each read the cards against the evidence and the code against its claims. They caught real problems: a stray input column that could override the tree's score, a card that understated what the fake-job fingerprints reveal, and a leaderboard figure attributed to the wrong source. All are fixed or disclosed.

The files you download are the files we reviewed: each repository was frozen after review and checked again, file by file, after upload.

Limits worth knowing

  • These are benchmark models. Each is built for one public dataset and its quirks. None is a production fraud system, and none should decide anything about a person without review.
  • The data is old or synthetic. Job ads from 2012 to 2014, a 2022 blocklist with random negatives, a simulator, and two days of 2013 card payments. Fraud moves; retrain on your own data before relying on any of them.
  • The released outputs were chosen after the test scores were in. Every candidate is printed, so you can judge how much that matters.
  • Fairness was not evaluated. The benchmark data carries no protected attributes to audit. Monitor error rates on your own population.

Try them

Each model card has the commands: download the repository, install its pinned requirements, and score the benchmark's test split in three lines of Python, or run eval/reproduce_fdb.py to check our numbers yourself. On a CPU, System One takes from 0.07 seconds a row on IP addresses to 0.53 seconds on job postings, which carry the longest evidence.

If you have real fraud logs, treat these as the recipe rather than the ceiling: the same Model Studio run trains a System One model on your own decisions, inside your own boundary.

All news

Read the cards, then bring your own data.

Talk to STEAVExplore CID