FinReasoning (Data Alignment) - Verification: leaderboard

Metric: Success-weighted score (0-100): the success rate (items whose query ran and whose answer came back in the required format) times the mean of answer accuracy, F1 over the cited data IDs and F1 over the cited fields, on the 600 verification items (whether a perturbed statement matches the database) of FinReasoning's Data Alignment track (answers checked against a structured A-share database by rule); Chinese-language tasks built from A-share research reports, news and market data (2023 to 2025); zero-shot at temperature 0.1; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 19 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3 Pro94.3#77
2Seed 1.893.5#136
3Qwen 3 Max92.3#201
4DeepSeek V391.8#312
5Qwen 3 32B91.6#424
6GPT-4o90.4#333
7Kimi K290.2#236
8Qwen 3 235B A22B90.1#304
9GLM-4.689.4#246
10Intern-S188.8#278
11GPT-588.3#91
12Llama 3.1 70B88.2#578
13Claude Sonnet 4.585.3#138
14Qwen 3 8B83.5#667
15Llama 3.1 8B78.4#1139

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=finreasoning-data-alignment-verification · How It Works · Data refreshed daily, snapshot 2026-10-11.