FinMTM - Self-Correction: leaderboard

Metric: FinMTM open-ended self-correction (L3: adversarial robustness and logical consistency) score (0-100) over 1,210 multi-turn financial sessions (about 4.25 turns each): half the mean LLM-judge rating of each turn on visual precision, financial logic, data accuracy, cross-modal verification and temporal awareness, half a session-level checklist for the task type; higher is better. Source: arxiv.org. Saturation forecast: Around December 2027. 22 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3 Pro58.8#77
2GPT-556.9#91
3Gemini 3 Flash55.4#93
4Qwen 3 VL 235B A22B Instruct54.5#264
5O352.8#121
6Qwen 3 VL 235B A22B (Thinking)52.5#228 (Qwen 3 VL 235B A22B)
7GLM-4.5V51.1#339
8Qwen 3 VL 32B Instruct50.8#276
9GPT-4o46.2#333
10Qwen 3 VL 30B A3B (Thinking)44.2#338 (Qwen 3 VL 30B A3B)
11InternVL3-78B43.6#345
12Qwen 3 VL 32B (Thinking)43.5#287 (Qwen 3 VL 32B)
13Qwen 2.5 VL 7B Instruct43.1#643
14Qwen 3 VL 30B A3B Instruct42.5#365
15Qwen 3 VL 4B (Thinking)42.5#471 (Qwen 3 VL 4B)

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=finmtm-self-correction · How It Works · Data refreshed daily, snapshot 2026-10-11.