EuroEval English: leaderboard
English LLM evaluation across 7 tasks: sentiment, NER, linguistic acceptability, reading comprehension, summarization, knowledge, and common-sense reasoning.
Metric: Average Score (%). Source: euroeval.com. Status: saturated. 343 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Llama 3.1 405B | 79.72 |
| 2 | GPT-5 (High) | 79.09 |
| 3 | Qwen 3 235B A22B 2507 Instruct | 78.94 |
| 4 | Llama 3.1 405B Instruct FP8 | 77.88 |
| 5 | Qwen 3 Next 80B A3B Instruct | 77.61 |
| 6 | O3 (2025-04-16) | 77.37 |
| 7 | Magistral Small 1.2 | 77.21 |
| 8 | GPT-5 | 77.16 |
| 9 | Yi 1.5 34B | 77.14 |
| 10 | Claude 3.7 Sonnet (20250219) (Thinking) | 76.61 |
| 11 | Qwen 3 235B A22B | 76.45 |
| 12 | GPT-5 (Minimal) | 76.06 |
| 13 | Mistral Small 3 | 76.05 |
| 14 | Gemini 2.5 Pro | 76 |
| 15 | Qwen 3 Next 80B A3B (Thinking) | 75.78 |
Interactive version: theaggregate.ai/benchmark?slug=euroeval-english · How It Works · Data refreshed daily, snapshot 2026-09-05.