Vals Multimodal Index — leaderboard

Benchmark consisting of a weighted performance across finance, coding, and education tasks. Showing the potential impact that LLM's can have on the economy.

Metric: Score (self-reported). Source: benchmarklist.com. Status: saturation imminent. 20 models tracked.

Top models

#ModelScore
1Claude Fable 5 (Max)74.15
2Claude Opus 4.8 (Max)70.89
3GPT-5.5 (xHigh)68.07
4Claude Opus 4.7 (High)67.36
5Gemini 3.5 Flash (High)62.93
6Claude Sonnet 4.6 (Max)60.57
7MiniMax-M359.97
8Gemini 3.1 Pro (Preview) (High)56.07
9GPT-5.4 Mini (xHigh)54.22
10Qwen 3.7 Plus53.89
11Gemini 3 Flash (Preview) (High)52.19
12Qwen 3.6 Plus51.52
13GPT-5.4 Nano (High)47.64
14Grok 4.3 (High)43.29
15Claude Haiku 4.542.88

Interactive version: theaggregate.ai/benchmark?slug=vals-multimodal-index · How the rankings work · Data refreshed daily, snapshot 2026-07-22.