EuroEval English — leaderboard
Comprehensive English LLM evaluation across 7 tasks: sentiment, NER, linguistic acceptability, reading comprehension, summarization, knowledge, and common-sense reasoning.
Metric: Average Score (%). Source: euroeval.com. Status: saturated. 343 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Llama 3.1 405B | 79.72 |
| 2 | GPT-5 (High) | 79.09 |
| 3 | Qwen 3 235B A22B 2507 Instruct | 78.94 |
| 4 | Llama 3.1 405B Instruct FP8 | 77.88 |
| 5 | Qwen 3 Next 80B A3B Instruct | 77.61 |
| 6 | O3 (2025-04-16) | 77.37 |
| 7 | Magistral Small | 77.21 |
| 8 | GPT-5 | 77.16 |
| 9 | Yi 1.5 34B | 77.14 |
| 10 | Claude 3.7 Sonnet (20250219) (Thinking) | 76.61 |
| 11 | Qwen 3 235B A22B | 76.45 |
| 12 | Gemini 3 Flash (Preview) (Non-reasoning) | 76.36 |
| 13 | GPT-5 (Minimal) | 76.06 |
| 14 | Mistral Small 3 | 76.05 |
| 15 | Gemini 2.5 Pro | 76 |
Interactive version: theaggregate.ai/benchmark?slug=euroeval-english · How the rankings work · Data refreshed daily, snapshot 2026-07-22.