FinFIRST - Raw-Information Acquisition: leaderboard

Metric: Weighted rubric score (%; the 7,322 rubric points for retrieving the correct values for the required entity, period, unit, definition and data version; 123 expert-authored bilingual financial research tasks, each scored against expert atomic rubric criteria (701 in all, weights summing to 100 per task) by a GLM-5.1 judge validated against finance experts (Cohen kappa 0.816); all models run in the same ReAct harness with web search, page visit and Python tools, temperature 1.0, one run; missing outputs count as failed). Source: arxiv.org. Saturation forecast: Around December 2026. 15 models tracked.

Top models

#ModelScore
1Claude Opus 588.98
2GPT-5.6 Sol86.85
3Kimi K384.84
4Qwen 3.8 Flash84.29
5GLM-5.383.11
6Qwen 3.8 27B81.49
7Gemini 3.7 Flash81.19
8Qwen 3.8 Max80.33
9DeepSeek V4 Pro (0813)80.09
10GLM-5.3 Flash79.5
11Ling-3.0-flash-fin78.43
12DeepSeek V4 Flash (0731)77.49
13GLM-5.269.75
14MiniMax-M365.51

Interactive version: theaggregate.ai/benchmark?slug=finfirst-raw-information-acquisition · How It Works · Data refreshed daily, snapshot 2026-09-26.