FinMTM - Multiple-Choice: leaderboard

Metric: FinMTM multiple-choice score (%) on the 1,982 negative-selection multiple-choice questions (select every incorrect option) built from expert-validated financial chart and report questions (Chinese and English), set-overlap scoring with no over-selection: any wrong option scores 0, otherwise credit is the share of correct options chosen; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 22 models tracked.

Top models

#ModelScoreOverall rank
1GPT-579.6#91
2Gemini 3 Pro78.4#77
3Gemini 3 Flash78.1#93
4O373.3#121
5GLM-4.5V51#339
6Qwen 3 VL 30B A3B (Thinking)49.4#338 (Qwen 3 VL 30B A3B)
7GPT-4o49.1#333
8Qwen 3 VL 235B A22B Instruct48.5#264
9Qwen 3 VL 30B A3B Instruct47.3#365
10Grok 4 Fast (Non-reasoning)46.8#242 (Grok 4 Fast)
11Qwen 3 VL 32B (Thinking)46.5#287 (Qwen 3 VL 32B)
12InternVL3-78B42.4#345
13Qwen 3 VL 235B A22B (Thinking)42.3#228 (Qwen 3 VL 235B A22B)
14Qwen 3 VL 32B Instruct39.9#276
15Qwen 3 VL 4B Instruct34.2#506

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=finmtm-multiple-choice · How It Works · Data refreshed daily, snapshot 2026-10-11.