FinReasoning (Semantic Consistency): leaderboard
Metric: Composite score (0-100): the mean of sentence-level F1 for locating the injected errors, an error-explanation score and a corrected-content score, each of the last two the mean of BERTScore, SimCSE similarity and a judge score fused from DeepSeek-V3 and Qwen3-235B against the expert-verified annotation, averaged over the Terminology, Fact and Logic categories (600 items each) of FinReasoning's Semantic Consistency track (long financial passages with one to three injected errors to locate, explain and correct); Chinese-language tasks built from A-share research reports, news and market data (2023 to 2025); zero-shot at temperature 0.1; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 19 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Seed 1.8 | 67.2 | #136 |
| 2 | GLM-4.6 | 65 | #246 |
| 3 | GPT-5 | 64.4 | #91 |
| 4 | Gemini 3 Pro | 61.5 | #77 |
| 5 | Qwen 3 Max | 59.4 | #201 |
| 6 | Claude Sonnet 4.5 | 58.6 | #138 |
| 7 | Kimi K2 | 58.2 | #236 |
| 8 | DeepSeek R1 | 55 | #245 |
| 9 | DeepSeek V3 | 52.9 | #312 |
| 10 | Intern-S1 | 44 | #278 |
| 11 | Qwen 3 235B A22B | 43.7 | #304 |
| 12 | GPT-4o | 42.4 | #333 |
| 13 | Qwen 3 32B | 42.4 | #424 |
| 14 | Qwen 3 8B | 38.7 | #667 |
| 15 | Llama 3.1 70B | 23 | #578 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=finreasoning-semantic-consistency · How It Works · Data refreshed daily, snapshot 2026-10-11.