CryptoAnalystBench - Fabricated Claims: leaderboard

Metric: Fabricated claims (%): share of the factual claims extracted from each answer that no retrieved tool output supports, directly or by simple derivation, on CryptoAnalystBench's 198 production crypto and DeFi analyst queries in 11 categories, one long-form answer per model from the authors' ReAct agent harness with web search and crypto data APIs; DeepSeek V3.1 extracts and verifies the claims; lower is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 5 models tracked.

Top models

#ModelScoreOverall rank
1GLM-4.71.22#185
2Kimi K2.51.51#139
3GPT-5.21.55#105
4GPT-OSS-120B5.49#330

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=cryptoanalystbench-fabricated-claims · How It Works · Data refreshed daily, snapshot 2026-10-11.