16,140 transactions replayed: Every newer AI model version caught less fraud at Coinbase

An internal benchmark run by Coinbase itself found that newer versions of three major AI model families caught fewer fraud transactions and a smaller share of the fraud value than their predecessors — even though the decision policy was…

Illustration: Thousands of black tokens slide down a steel chute, with a worn mesh sieve catching only a few red tokens as the rest slip past.
Illustration

16,140 transactions replayed: Every newer AI model version caught less fraud at Coinbase

An internal benchmark run by Coinbase itself found that newer versions of three major AI model families caught fewer fraud transactions and a smaller share of the fraud value than their predecessors — even though the decision policy was unchanged. The findings were published on October 7, 2026 and reported by CryptoSlate on October 10, and they challenge a widespread industry assumption: that upgrading to the latest frontier model automatically improves the production systems it is dropped into.

What was measured

The evaluation was a historical replay of fraud screening in Coinbase's Onramp service, where cryptocurrency is purchased with traditional payment methods. According to CryptoSlate's write-up of Coinbase's published results, 16,140 transactions from 7,293 users were replayed, including 813 confirmed fraud transactions. The cohort covered nine weeks before Coinbase's risk agent was rolled out in production, and the setup retained all matured fraud trajectories while legitimate traffic was sampled.

This allowed Coinbase to run the same historical traffic through different model versions and compare them under identical conditions — the same decision policy and the same fraud ground truth, built from the traffic that actually occurred before the risk agent went live.

The models and the numbers

Coinbase compared three model pairs: Opus 4.5 against Opus 5, Sonnet 4.6 against Sonnet 5, and GPT-5.4 against GPT-5.6 (sol). According to the figures reported by CryptoSlate, every single newer version had lower recall (the share of actual fraud that gets caught), lower F1 (a combined precision and recall score), and lower dollar-weighted recall (the share of fraud value that gets detected).

The spread between the pairs was substantial:

  • Sonnet saw the largest drop: recall fell 22.2 percentage points, and dollar-weighted recall fell 22.9 points. Precision also got worse — a smaller share of the transactions the model flagged as fraud were actually fraud.
  • Opus showed a much smaller but still real decline: recall dropped 0.8 points, with lower precision for the newer version as well.
  • GPT revealed a different pattern: precision rose 11.5 percentage points, but recall fell 20.7 points and dollar-weighted recall fell 21.8 points. The flags became more accurate — while more fraud trajectories and more fraud value slipped through the replay undetected.

Why one number can hide a regression

The GPT case is the core of what kind of error this type of evaluation surfaces. A team monitoring a single metric — say, precision, which measures how "correct" the fraud flags are — could see an improvement of more than eleven points and conclude the upgrade was a success. But precision and recall point in opposite directions when a model becomes more conservative: it flags fewer transactions, and among those it flags, more are right. The result is cleaner flags and more fraud getting through.

For a production screener in a payments system, it is recall — not precision alone — that determines how much fraud is actually stopped. That is why dollar-weighted recall may be the most telling metric in the eval: it weights every missed recall by the transaction's value, and on that measure all three newer model versions fell.

The operational message is not that new models are worse in any general sense, but that a model upgrade changes the model's behavior in ways that can shift the balance between false positives and missed fraud trajectories — even when the policy, thresholds, and rules around the model are entirely unchanged. Coinbase's eval illustrates the value of replaying historical traffic through the new model before rollout, with the metric chain intact, rather than updating and hoping for the best.

Caveats and open questions

The findings have clear limitations, and most of them are stacked in the source material itself.

First, the replay does not establish real customer losses. The replay measures model behavior in a specific test configuration — nine weeks of historical traffic, fixed policy, sampled legitimate traffic — not what actually happened to Coinbase's users in production. CryptoSlate states this caveat explicitly.

Second, every number in this story rests on a single secondary source. Coinbase's own blog post or eval report is not available in the source material here, so methodology details, model version names, and all delta figures come from CryptoSlate's write-up and cannot be verified against the primary source. Version labels such as "Opus 5" and "GPT-5.6 (sol)" have not been cross-checked against any official release in the available material.

Third, Coinbase's stated remediation is unknown. The source text is cut off mid-sentence — according to CryptoSlate, Coinbase said it could identify the regressions with, but the sentence stops there — so what the company did with the findings, which models were actually rolled out, and whether the policy was later adjusted remain open questions.

It is also worth noting what the eval does not say: it says nothing about why the newer models performed worse, whether the pattern holds outside this dataset, or whether other model families would show the same thing. What it documents — if the numbers hold — is something sharper and more useful: that a frontier model upgrade can make a production fraud filter worse without any change to the decision logic, and that only a full replay evaluation with the right metrics surfaces it before customers do.

AIMag.no
AIMag.no
The AIMag.no editorial team covers artificial intelligence, tools, research, and regulation.

Get the best of AI MAG in your inbox

News, analysis, and ideas at the intersection of AI and society.