Loading prices…
〽️NEUTRAL

AI Models Miss More Payment Fraud in Coinbase Test

Coinbase’s replay shows that model upgrades can weaken fraud coverage under a fixed screening policy, while a specialized Qwen3.5-9B model improved detection and cut latency.

Coinbase’s replay of 16,140 transactions across 7,293 users found that newer versions of three major AI model families caught less payment fraud than their predecessors. The test included 813 confirmed fraudulent transactions and held the decision policy constant. Sonnet’s recall fell 22.2 percentage points, while GPT’s recall dropped 20.7 points and its dollar-weighted recall declined 21.8 points.

Why it matters

The results challenge the assumption that a newer general-purpose model automatically improves an existing payment screener. Recall measures the share of fraud cases caught, while dollar-weighted recall measures the share of total fraud value detected. GPT delivered a higher precision score, up 11.5 points, but its flags were more accurate as a group while more fraud cases and value escaped detection.

The historical replay isolates model behavior within one screening setup. It does not establish customer losses from deploying the tested versions, and Coinbase said it could identify the regressions without determining their cause. The proprietary dataset also limits independent replication.

Market impact

Coinbase’s separate evaluation of a post-trained Qwen3.5-9B model points to a different path. Specialized with historical fraud outcomes and deterministic rewards, it exceeded Opus 4.5 across four fraud-detection metrics. F1 improved 9.6 percentage points and dollar-weighted recall rose 35.4 points.

Production measurements also put median end-to-end request latency at 0.683 seconds for the custom model, versus 1.515 seconds for Opus 4.5, a 55% relative reduction. Coinbase recommends testing each candidate inside the actual decision setup, then assessing prompts, thresholds, latency, reliability and cost separately.

Frequently asked questions

  1. How many transactions did Coinbase test in its payment-fraud replay?

    The replay covered 16,140 transactions across 7,293 users, including 813 confirmed fraudulent transactions.

  2. Which AI model showed the largest recall decline in Coinbase’s test?

    Sonnet’s recall fell 22.2 percentage points. GPT’s recall declined 20.7 points, while Opus’s recall fell 0.8 points.

  3. Why was GPT’s higher precision not enough to improve fraud screening?

    GPT’s precision rose 11.5 percentage points, but its recall and dollar-weighted recall fell. Its flags were more accurate, while more fraud cases and value escaped detection.

  4. How did the specialized Qwen3.5-9B model perform?

    In a separate evaluation, Qwen3.5-9B exceeded Opus 4.5 across four fraud-detection metrics. F1 improved 9.6 points and dollar-weighted recall rose 35.4 points.

  5. What did Coinbase recommend before upgrading a payment-screening model?

    Coinbase recommended testing the candidate inside the actual decision setup, then evaluating prompts and thresholds separately alongside latency, reliability and cost.

Source attribution
Aggregated from CryptoSlate · Verified · Last refreshed 1h ago
Open original →