Skip to content

The Aggregate

Unified LLM rankings, daily

Models What's New Trends Benchmarks Skill Maps
How It Works Guesswork Metrics
The Aggregate
Loading data...

AA Global-MMLU-Lite - English — leaderboard

Metric: Accuracy (%). Source: artificialanalysis.ai. 120 models tracked.

Top models

#ModelScore
1Claude Opus 4.6 (Adaptive Reasoning, Max Effort)95.17
2Gemini 3.1 Pro (Preview)94.75
3Claude Opus 4.5 (Non-reasoning)94.17
4GPT-5.1 (High)94.08
5GPT-5 (High)93.75
6Claude Opus 4.5 (Thinking)93.75
7GPT-5 (Medium)93.67
8Gemini 3 Pro (Preview) (High)93.67
9Claude Sonnet 4.5 (Thinking)93.5
10Grok 4.20 0309 (Reasoning)93.42
11Gemini 2.5 Pro93.33
12DeepSeek V3.2 Speciale93.25
13Grok 492.92
14Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort)92.92
15Qwen 3.5 122B A10B (Reasoning)92.83

Interactive version: theaggregate.ai/benchmark?slug=aa-global-mmlu-lite-english · How the rankings work · Data refreshed daily, snapshot 2026-07-22.

Built by Mikhail Doroshenko — AI researcher, co-author of Humanity’s Last Exam. Independent project; no affiliation with any AI lab.