FinMTM - Single-Choice: leaderboard

Metric: FinMTM single-choice score (%) on the 1,982 single-choice questions (distractors written by Gemini 3 Pro) built from expert-validated financial chart and report questions (Chinese and English), set-overlap scoring with no over-selection: any wrong option scores 0, otherwise credit is the share of correct options chosen; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 22 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3 Pro92.1#77
2Gemini 3 Flash91.9#93
3GPT-589#91
4O385.8#121
5Qwen 3 VL 32B Instruct84.5#276
6Qwen 3 VL 32B (Thinking)83.4#287 (Qwen 3 VL 32B)
7Qwen 3 VL 235B A22B Instruct81.3#264
8Qwen 3 VL 235B A22B (Thinking)80.5#228 (Qwen 3 VL 235B A22B)
9GPT-4o79.3#333
10Qwen 3 VL 30B A3B Instruct77.2#365
11InternVL3-78B75.6#345
12GLM-4.5V73.7#339
13Qwen 2.5 VL 7B Instruct73.4#643
14Qwen 3 VL 4B Instruct73.3#506
15Qwen 3 VL 30B A3B (Thinking)71.5#338 (Qwen 3 VL 30B A3B)

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=finmtm-single-choice · How It Works · Data refreshed daily, snapshot 2026-10-11.