FinMTM - Financial Agent: leaderboard

Metric: FinMTM financial agent score (0-100) over 1,000 tool-use tasks with MCP financial tools: tool-call quality (recall-weighted F2 of predicted against reference calls, up to 25) plus judged reasoning (up to 25) and answer correctness against the ground truth (up to 50); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 22 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3 Flash62.6#93
2Gemini 3 Pro54.3#77
3GPT-549.7#91
4Qwen 3 VL 235B A22B (Thinking)41.5#228 (Qwen 3 VL 235B A22B)
5Grok 4 Fast (Non-reasoning)39.7#242 (Grok 4 Fast)
6Qwen 3 VL 235B A22B Instruct38.7#264
7O335.2#121
8GPT-4o34.8#333
9GLM-4.5V32.4#339
10Qwen 3 VL 32B (Thinking)28.6#287 (Qwen 3 VL 32B)
11Qwen 3 VL 32B Instruct25.1#276
12Qwen 3 VL 30B A3B (Thinking)23.3#338 (Qwen 3 VL 30B A3B)
13InternVL3-78B22.8#345
14Qwen 3 VL 30B A3B Instruct20.8#365
15Qwen 3 VL 4B Instruct19.1#506

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=finmtm-financial-agent · How It Works · Data refreshed daily, snapshot 2026-10-11.