Vals Index — leaderboard

Benchmark consisting of a weighted performance across finance and coding tasks. Showing the potential impact that LLM's can have on the economy.

Metric: Score (self-reported). Source: benchmarklist.com. Status: saturation imminent. 25 models tracked.

Top models

#ModelScore
1Claude Fable 5 (Max)75.14
2Claude Opus 4.8 (Max)70.36
3GPT-5.5 (xHigh)67.95
4Claude Opus 4.7 (High)66.1
5Gemini 3.5 Flash (High)62.75
6Claude Sonnet 4.6 (Max)60.06
7MiniMax-M358.94
8Qwen 3.7 Max (Max)57.49
9DeepSeek V4 Pro (Max)55.62
10Gemini 3.1 Pro (Preview) (High)53.77
11GLM-5.152.45
12GPT-5.4 Mini (xHigh)52.42
13Qwen 3.7 Plus52.33
14Gemini 3 Flash (Preview) (High)49.55
15Qwen 3.6 Plus48.89

Interactive version: theaggregate.ai/benchmark?slug=vals-index · How the rankings work · Data refreshed daily, snapshot 2026-07-22.