FinMTM - Financial Agent (Fuzzed): leaderboard

Metric: FinMTM financial agent score (0-100) over 1,000 tool-use tasks with MCP financial tools whose queries obfuscate entity names through knowledge-base substitutions and masked image regions: tool-call quality (recall-weighted F2 of predicted against reference calls, up to 25) plus judged reasoning (up to 25) and answer correctness against the ground truth (up to 50); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 22 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3 Flash53.6#93
2Gemini 3 Pro48.3#77
3GPT-535.9#91
4Qwen 3 VL 235B A22B (Thinking)35.2#228 (Qwen 3 VL 235B A22B)
5Qwen 3 VL 235B A22B Instruct32.1#264
6O331.4#121
7Grok 4 Fast (Non-reasoning)30.2#242 (Grok 4 Fast)
8GPT-4o29.7#333
9GLM-4.5V26.5#339
10Qwen 3 VL 32B (Thinking)23.2#287 (Qwen 3 VL 32B)
11Qwen 3 VL 32B Instruct19.6#276
12Qwen 3 VL 30B A3B (Thinking)18.9#338 (Qwen 3 VL 30B A3B)
13InternVL3-78B18.2#345
14Qwen 3 VL 30B A3B Instruct16.2#365
15Qwen 3 VL 4B Instruct15.1#506

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=finmtm-financial-agent-fuzzed · How It Works · Data refreshed daily, snapshot 2026-10-11.