CorpusQA (128K) - Financial (English): leaderboard

Metric: Accuracy (%; 83 questions over English financial filings, 128K-token corpora, DeepSeek-V3 judge). Source: arxiv.org. Saturation forecast: Estimated already saturated. 8 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro80.72
2GPT-578.31
3GPT-5 Mini78.31
4Gemini 2.5 Flash74.7
5Qwen 3 235B A22B 2507 (Thinking)74.7
6DeepSeek R1 052873.49
7MiniMax M1 80k67.07
8Qwen 3 30B A3B 2507 (Thinking)61.45

Interactive version: theaggregate.ai/benchmark?slug=corpusqa-128k-financial-english · How It Works · Data refreshed daily, snapshot 2026-09-25.