FinFIRST: leaderboard

Metric: Weighted rubric score (%; Loose Pass, share of the 12,300 expert-weighted rubric points satisfied; 123 expert-authored bilingual financial research tasks, each scored against expert atomic rubric criteria (701 in all, weights summing to 100 per task) by a GLM-5.1 judge validated against finance experts (Cohen kappa 0.816); all models run in the same ReAct harness with web search, page visit and Python tools, temperature 1.0, one run; missing outputs count as failed). Source: arxiv.org. Saturation forecast: Around December 2026. 15 models tracked.

Top models

#ModelScore
1Claude Opus 587.61
2GPT-5.6 Sol85.92
3Qwen 3.8 Flash81.23
4Kimi K380.83
5GLM-5.380.61
6Qwen 3.8 Max77.4
7Qwen 3.8 27B77.28
8Gemini 3.7 Flash76.75
9GLM-5.3 Flash76.64
10DeepSeek V4 Pro (0813)75.41
11Ling-3.0-flash-fin75.07
12DeepSeek V4 Flash (0731)71.37
13GLM-5.265.89
14MiniMax-M360.7

Interactive version: theaggregate.ai/benchmark?slug=finfirst · How It Works · Data refreshed daily, snapshot 2026-09-26.