FinFIRST - Source Verification: leaderboard

Metric: Weighted rubric score (%; the 2,085 rubric points for using authoritative, relevant and time-valid sources; 123 expert-authored bilingual financial research tasks, each scored against expert atomic rubric criteria (701 in all, weights summing to 100 per task) by a GLM-5.1 judge validated against finance experts (Cohen kappa 0.816); all models run in the same ReAct harness with web search, page visit and Python tools, temperature 1.0, one run; missing outputs count as failed). Source: arxiv.org. Saturation forecast: Around December 2026. 15 models tracked.

Top models

#ModelScore
1Claude Opus 589.59
2GPT-5.6 Sol88.68
3Qwen 3.8 Flash83.36
4Ling-3.0-flash-fin82.45
5GLM-5.3 Flash80.77
6GLM-5.378.56
7Qwen 3.8 Max77.75
8Qwen 3.8 27B76.83
9DeepSeek V4 Pro (0813)73.57
10DeepSeek V4 Flash (0731)71.65
11Kimi K370.7
12GLM-5.268.3
13Gemini 3.7 Flash66.62
14MiniMax-M362.97

Interactive version: theaggregate.ai/benchmark?slug=finfirst-source-verification · How It Works · Data refreshed daily, snapshot 2026-09-26.