FinReasoning (Semantic Consistency) - Logic: leaderboard

Metric: Composite score (0-100): the mean of sentence-level F1 for locating the injected errors, an error-explanation score and a corrected-content score, each of the last two the mean of BERTScore, SimCSE similarity and a judge score fused from DeepSeek-V3 and Qwen3-235B against the expert-verified annotation, on the 600 logic-error items (reasoning-chain, discourse-relation and context-inconsistency errors) of FinReasoning's Semantic Consistency track (long financial passages with one to three injected errors to locate, explain and correct); Chinese-language tasks built from A-share research reports, news and market data (2023 to 2025); zero-shot at temperature 0.1; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 19 models tracked.

Top models

#ModelScoreOverall rank
1Seed 1.863.3#136
2GLM-4.662.5#246
3Kimi K259.1#236
4GPT-557.6#91
5Qwen 3 Max57.1#201
6DeepSeek V354#312
7Claude Sonnet 4.551.8#138
8Gemini 3 Pro50.2#77
9DeepSeek R146.7#245
10GPT-4o37.6#333
11Qwen 3 8B37.6#667
12Intern-S135.4#278
13Qwen 3 32B34.2#424
14Qwen 3 235B A22B33#304
15Llama 3.1 70B20.4#578

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=finreasoning-semantic-consistency-logic · How It Works · Data refreshed daily, snapshot 2026-10-11.