FinReasoning (Data Alignment) - Reasoning: leaderboard

Metric: Success-weighted score (0-100): the success rate (items whose query ran and whose answer came back in the required format) times the mean of answer accuracy, F1 over the cited data IDs and F1 over the cited fields, on the 400 rule-driven reasoning items (61 executable financial-analysis rules) of FinReasoning's Data Alignment track (answers checked against a structured A-share database by rule); Chinese-language tasks built from A-share research reports, news and market data (2023 to 2025); zero-shot at temperature 0.1; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 19 models tracked.

Top models

#ModelScoreOverall rank
1GPT-568.5#91
2Qwen 3 Max66.3#201
3Qwen 3 32B65#424
4Kimi K263.2#236
5GPT-4o63.1#333
6Claude Sonnet 4.562.7#138
7Llama 3.1 70B61.3#578
8DeepSeek V360.4#312
9GLM-4.658.4#246
10Gemini 3 Pro56.8#77
11Seed 1.855.6#136
12Qwen 3 8B54.2#667
13Qwen 3 235B A22B54.2#304
14Intern-S153.9#278
15Llama 3.1 8B42.6#1139

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=finreasoning-data-alignment-reasoning · How It Works · Data refreshed daily, snapshot 2026-10-11.