FinReasoning (Semantic Consistency) - Fact: leaderboard

Metric: Composite score (0-100): the mean of sentence-level F1 for locating the injected errors, an error-explanation score and a corrected-content score, each of the last two the mean of BERTScore, SimCSE similarity and a judge score fused from DeepSeek-V3 and Qwen3-235B against the expert-verified annotation, on the 600 factual-error items (relation, entity and context errors) of FinReasoning's Semantic Consistency track (long financial passages with one to three injected errors to locate, explain and correct); Chinese-language tasks built from A-share research reports, news and market data (2023 to 2025); zero-shot at temperature 0.1; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 19 models tracked.

Top models

#ModelScoreOverall rank
1Seed 1.874.4#136
2GLM-4.673.4#246
3GPT-571.4#91
4Gemini 3 Pro71.1#77
5Qwen 3 Max70#201
6Claude Sonnet 4.569#138
7DeepSeek R166.9#245
8Kimi K266#236
9DeepSeek V363.1#312
10Qwen 3 235B A22B57.7#304
11Intern-S157.1#278
12Qwen 3 32B54.9#424
13GPT-4o53.8#333
14Qwen 3 8B46.1#667
15Llama 3.1 70B28.8#578

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=finreasoning-semantic-consistency-fact · How It Works · Data refreshed daily, snapshot 2026-10-11.