FinReasoning (Semantic Consistency): leaderboard

Metric: Composite score (0-100): the mean of sentence-level F1 for locating the injected errors, an error-explanation score and a corrected-content score, each of the last two the mean of BERTScore, SimCSE similarity and a judge score fused from DeepSeek-V3 and Qwen3-235B against the expert-verified annotation, averaged over the Terminology, Fact and Logic categories (600 items each) of FinReasoning's Semantic Consistency track (long financial passages with one to three injected errors to locate, explain and correct); Chinese-language tasks built from A-share research reports, news and market data (2023 to 2025); zero-shot at temperature 0.1; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 19 models tracked.

Top models

#ModelScoreOverall rank
1Seed 1.867.2#136
2GLM-4.665#246
3GPT-564.4#91
4Gemini 3 Pro61.5#77
5Qwen 3 Max59.4#201
6Claude Sonnet 4.558.6#138
7Kimi K258.2#236
8DeepSeek R155#245
9DeepSeek V352.9#312
10Intern-S144#278
11Qwen 3 235B A22B43.7#304
12GPT-4o42.4#333
13Qwen 3 32B42.4#424
14Qwen 3 8B38.7#667
15Llama 3.1 70B23#578

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=finreasoning-semantic-consistency · How It Works · Data refreshed daily, snapshot 2026-10-11.