FinMTM - Calculation: leaderboard

Metric: FinMTM open-ended calculation (L2: multi-step numerical calculation and chart value estimation) score (0-100) over 1,893 multi-turn financial sessions (about 4.25 turns each): half the mean LLM-judge rating of each turn on visual precision, financial logic, data accuracy, cross-modal verification and temporal awareness, half a session-level checklist for the task type; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 22 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3 Pro82.8#77
2Qwen 3 VL 235B A22B Instruct80.9#264
3GPT-580.7#91
4Qwen 3 VL 32B Instruct80.7#276
5GLM-4.5V79.6#339
6Qwen 3 VL 235B A22B (Thinking)79.4#228 (Qwen 3 VL 235B A22B)
7O378.6#121
8InternVL3-78B77.6#345
9GPT-4o76.8#333
10Qwen 3 VL 30B A3B Instruct76.5#365
11Gemini 3 Flash76#93
12Qwen 2.5 VL 7B Instruct73.4#643
13Qwen 3 VL 4B Instruct71.2#506
14Qwen 3 VL 32B (Thinking)68.6#287 (Qwen 3 VL 32B)
15Qwen 3 VL 4B (Thinking)68.5#471 (Qwen 3 VL 4B)

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=finmtm-calculation · How It Works · Data refreshed daily, snapshot 2026-10-11.