FinReasoning (Data Alignment) - Calculation: leaderboard

Metric: Success-weighted score (0-100): the success rate (items whose query ran and whose answer came back in the required format) times the mean of answer accuracy, F1 over the cited data IDs and F1 over the cited fields, on the 800 calculation items (35 indicator formulas across fields, dates and companies) of FinReasoning's Data Alignment track (answers checked against a structured A-share database by rule); Chinese-language tasks built from A-share research reports, news and market data (2023 to 2025); zero-shot at temperature 0.1; higher is better. Source: arxiv.org. Saturation forecast: Around March 2027. 19 models tracked.

Top models

#ModelScoreOverall rank
1GPT-572.8#91
2Gemini 3 Pro69.9#77
3Kimi K269.3#236
4DeepSeek R168.9#245
5GLM-4.667.5#246
6Qwen 3 Max67.1#201
7Seed 1.866.6#136
8Claude Sonnet 4.564.6#138
9DeepSeek V363.7#312
10Qwen 3 235B A22B63.1#304
11Intern-S162.9#278
12Llama 3.1 70B61.3#578
13GPT-4o61#333
14Qwen 3 32B55.7#424
15Qwen 3 8B51.8#667

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=finreasoning-data-alignment-calculation · How It Works · Data refreshed daily, snapshot 2026-10-11.