TraderBench - Knowledge Retrieval: leaderboard

Metric: Knowledge Retrieval score (0-100) on 18 tasks drawn from BizFinBench and PRBench that require retrieving figures from SEC filings and market data servers (exact match plus LLM rubric), in the authors' two-agent setup (an A2A Candidate Agent with the same six financial MCP servers for every model; GPT-5.2 judge at temperature 0 for rubric-scored sections); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 12 models tracked.

Top models

#ModelScoreOverall rank
1Grok 4.1 Fast61.6#208
2Gemini 3 Pro52.4#77
3GPT-5.250.6#105
4GPT-OSS-20B42.4#499
5GLM-4.7 Flash38.9#496
6Kimi K2.537.1#139
7GPT-OSS-120B32.5#330
8GPT-4o32.3#333
9Qwen 3 32B (Non-reasoning)14.4#424 (Qwen 3 32B)
10Qwen 3 8B10.5#667
11Qwen 3 30B A3B9.4#488
12Gemma 3 27B9.3#596

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=traderbench-knowledge-retrieval · How It Works · Data refreshed daily, snapshot 2026-10-11.