FinMTM - Comprehension: leaderboard

Metric: FinMTM open-ended comprehension (L1: entity recognition and spatial awareness) score (0-100) over 2,082 multi-turn financial sessions (about 4.25 turns each): half the mean LLM-judge rating of each turn on visual precision, financial logic, data accuracy, cross-modal verification and temporal awareness, half a session-level checklist for the task type; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 22 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3 Pro87.5#77
2GPT-586.9#91
3Qwen 3 VL 235B A22B Instruct85.5#264
4GLM-4.5V85.4#339
5Qwen 3 VL 235B A22B (Thinking)84.5#228 (Qwen 3 VL 235B A22B)
6Qwen 3 VL 32B Instruct84.3#276
7O383.8#121
8Gemini 3 Flash82.2#93
9Qwen 3 VL 30B A3B Instruct82.1#365
10Qwen 3 VL 30B A3B (Thinking)80.7#338 (Qwen 3 VL 30B A3B)
11Qwen 3 VL 32B (Thinking)80.3#287 (Qwen 3 VL 32B)
12GPT-4o77.2#333
13InternVL3-78B76.2#345
14Qwen 3 VL 4B Instruct74.5#506
15Qwen 2.5 VL 7B Instruct74.3#643

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=finmtm-comprehension · How It Works · Data refreshed daily, snapshot 2026-10-11.