FinReasoning (Semantic Consistency) - Terminology: leaderboard

Metric: Composite score (0-100): the mean of sentence-level F1 for locating the injected errors, an error-explanation score and a corrected-content score, each of the last two the mean of BERTScore, SimCSE similarity and a judge score fused from DeepSeek-V3 and Qwen3-235B against the expert-verified annotation, on the 600 terminology-error items (improper usage, inconsistency, confusion) of FinReasoning's Semantic Consistency track (long financial passages with one to three injected errors to locate, explain and correct); Chinese-language tasks built from A-share research reports, news and market data (2023 to 2025); zero-shot at temperature 0.1; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 19 models tracked.

Top models

#ModelScoreOverall rank
1GPT-564.3#91
2Seed 1.863.9#136
3Gemini 3 Pro63.1#77
4GLM-4.659#246
5Claude Sonnet 4.554.9#138
6DeepSeek R151.6#245
7Qwen 3 Max51.2#201
8Kimi K249.5#236
9DeepSeek V341.7#312
10Qwen 3 235B A22B40.3#304
11Intern-S139.5#278
12Qwen 3 32B38.1#424
13GPT-4o35.7#333
14Qwen 3 8B32.4#667
15Llama 3.1 70B19.7#578

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=finreasoning-semantic-consistency-terminology · How It Works · Data refreshed daily, snapshot 2026-10-11.