FinFIRST - Computation and Answer Formation: leaderboard

Metric: Weighted rubric score (%; the 2,893 rubric points for calculation, reasoning and an evidence-supported final answer; 123 expert-authored bilingual financial research tasks, each scored against expert atomic rubric criteria (701 in all, weights summing to 100 per task) by a GLM-5.1 judge validated against finance experts (Cohen kappa 0.816); all models run in the same ReAct harness with web search, page visit and Python tools, temperature 1.0, one run; missing outputs count as failed). Source: arxiv.org. Saturation forecast: Around December 2026. 15 models tracked.

Top models

#ModelScore
1Claude Opus 582.72
2GPT-5.6 Sol81.58
3Kimi K377.98
4GLM-5.375.77
5Gemini 3.7 Flash72.8
6Qwen 3.8 Flash71.93
7Qwen 3.8 Max69.72
8Qwen 3.8 27B66.92
9GLM-5.3 Flash66.44
10DeepSeek V4 Pro (0813)64.88
11Ling-3.0-flash-fin61.25
12DeepSeek V4 Flash (0731)55.69
13GLM-5.254.37
14MiniMax-M346.87

Interactive version: theaggregate.ai/benchmark?slug=finfirst-computation-and-answer-formation · How It Works · Data refreshed daily, snapshot 2026-09-26.