MixEval — leaderboard

Derives queries from real-world user interactions and matches them with similar existing benchmark queries. Achieves 0.96 correlation with Chatbot Arena while being much cheaper to run.

Metric: Score. Source: mixeval.github.io. Status: saturation imminent. 52 models tracked.

Top models

#ModelScore
1O1 Preview72
2Llama 3.1 405B Instruct66.2
3GPT-4o (2024-05-13)64.7
4Claude 3 Opus63.5
5GPT-4 Turbo62.6
6Mistral Large 2 (Jul)57.4
7Yi Large (Preview)56.8
8Llama 3 70B Instruct55.9
9Claude 3 Sonnet54
10MAmmoTH2-8x7B-Plus51.8
11DeepSeek V251.7
12GPT-4o Mini51.6
13Command-R+51.4
14Yi 1.5 34B Chat51.2
15Mistral Large50.3

Interactive version: theaggregate.ai/benchmark?slug=mixeval · How the rankings work · Data refreshed daily, snapshot 2026-07-22.