FinReasoning (Data Alignment): leaderboard

Metric: Success-weighted score (0-100): the success rate (items whose query ran and whose answer came back in the required format) times the mean of answer accuracy, F1 over the cited data IDs and F1 over the cited fields, averaged over the Verification (600 items), Calculation (800) and Reasoning (400) categories of FinReasoning's Data Alignment track (answers checked against a structured A-share database by rule); Chinese-language tasks built from A-share research reports, news and market data (2023 to 2025); zero-shot at temperature 0.1; higher is better. Source: arxiv.org. Saturation forecast: Around April 2027. 19 models tracked.

Top models

#ModelScoreOverall rank
1GPT-576.5#91
2Qwen 3 Max75.2#201
3Kimi K274.2#236
4Gemini 3 Pro73.7#77
5DeepSeek V372#312
6Seed 1.871.9#136
7GPT-4o71.5#333
8Claude Sonnet 4.570.9#138
9Qwen 3 32B70.8#424
10Llama 3.1 70B70.3#578
11Qwen 3 235B A22B69.1#304
12Intern-S168.5#278
13Qwen 3 8B63.2#667
14Llama 3.1 8B56.7#1139
15DeepSeek R139.1#245

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=finreasoning-data-alignment · How It Works · Data refreshed daily, snapshot 2026-10-11.