EuroEval English — leaderboard

Comprehensive English LLM evaluation across 7 tasks: sentiment, NER, linguistic acceptability, reading comprehension, summarization, knowledge, and common-sense reasoning.

Metric: Average Score (%). Source: euroeval.com. Status: saturated. 343 models tracked.

Top models

#ModelScore
1Llama 3.1 405B79.72
2GPT-5 (High)79.09
3Qwen 3 235B A22B 2507 Instruct78.94
4Llama 3.1 405B Instruct FP877.88
5Qwen 3 Next 80B A3B Instruct77.61
6O3 (2025-04-16)77.37
7Magistral Small77.21
8GPT-577.16
9Yi 1.5 34B77.14
10Claude 3.7 Sonnet (20250219) (Thinking)76.61
11Qwen 3 235B A22B76.45
12Gemini 3 Flash (Preview) (Non-reasoning)76.36
13GPT-5 (Minimal)76.06
14Mistral Small 376.05
15Gemini 2.5 Pro76

Interactive version: theaggregate.ai/benchmark?slug=euroeval-english · How the rankings work · Data refreshed daily, snapshot 2026-07-22.