Market desk Bitcoin Ethereum Altcoins DeFi Stablecoins Markets & Trading

Coinbase Study: Newer AI Models Caught Fewer Fraudulent Payments in Payment Screening Test

Coinbase's historical analysis of its Onramp payment service found that newer versions of three major AI model families detected fewer fraudulent transactions and less fraud value despite an unchanged decision policy, challenging assumptions about model upgrades.
5 hours ago 14 views
Coinbase Study: Newer AI Models Caught Fewer Fraudulent Payments in Payment Screening Test

Coinbase reported that newer versions of three major AI model families caught fewer fraudulent payments in a historical test of its Onramp payment screening service. The findings suggest that upgrading to newer model versions does not automatically improve fraud detection performance.

The company evaluated the models using a fixed historical replay of 16,140 transactions across 7,293 users, including 813 confirmed fraudulent transactions. The dataset covered nine weeks before a risk agent rollout. Each candidate model reviewed recent transaction behavior under the same policy for converting risk classifications into decisions.

Results Across Model Families

Coinbase compared three model pairs: Opus 4.5 with Opus 5, Sonnet 4.6 with Sonnet 5, and GPT-5.4 with GPT-5.6. All newer versions showed lower recall, meaning they caught a smaller share of fraudulent cases. Every newer model also had lower dollar-weighted recall, which measures how much of the total fraud value the model detected.

Sonnet's recall declined 22.2 percentage points, while Opus's recall fell 0.8 points. Both Sonnet and Opus showed lower precision in their newer versions, indicating that a smaller share of transactions they flagged as fraudulent were actually fraudulent.

GPT's results illustrated how a single improved metric can mask broader performance decline. Its precision rose 11.5 percentage points in the newer version, but recall fell 20.7 points and dollar-weighted recall dropped 21.8 points. While fraud flags became more accurate, more fraudulent cases and value escaped detection.

Custom Model and Latency Testing

In separate testing, Coinbase reported that a custom model based on Qwen3.5-9B outperformed Opus 4.5 across four fraud-detection metrics. The F1 score improved 9.6 percentage points and dollar-weighted recall rose 35.4 points. The custom model was specialized using historical fraud outcomes and deterministic rewards.

Production measurements showed the custom model achieved 0.683 seconds median end-to-end latency compared to 1.515 seconds for Opus 4.5, a 55 percent relative reduction.

Implications for Payment Providers

Coinbase noted that neither the historical replay nor the latency tests measured actual customer losses from deploying the newer model versions. The company stated it identified the performance regressions without establishing their cause.

For payment providers, Coinbase recommends testing whether candidate models improve fraud coverage within their actual decision setup before evaluating other changes, while considering latency, reliability, and cost alongside detection quality.

Market snapshot

Top cryptocurrency prices

Explore all prices
BitcoinBTC $83,080.91+0.55% EthereumETH $2,508.65+0.77% Tether USDUSDT $0.9997-0.01% BNBBNB $750.96+1.16% XRPXRP $1.40+0.60% USDCUSDC $1.000.00% SolanaSOL $110.48+1.14% TRONTRX $0.3309-0.36% HyperliquidHYPE $85.47+1.35% ZcashZEC $1,223.69+0.92%
Prices by Coinranking. Informational only.