EuroEval English: leaderboard

English LLM evaluation across 7 tasks: sentiment, NER, linguistic acceptability, reading comprehension, summarization, knowledge, and common-sense reasoning.

Metric: Average Score (%). Source: euroeval.com. Status: saturated. 343 models tracked.

Top models

#ModelScore
1Llama 3.1 405B79.72
2GPT-5 (High)79.09
3Qwen 3 235B A22B 2507 Instruct78.94
4Llama 3.1 405B Instruct FP877.88
5Qwen 3 Next 80B A3B Instruct77.61
6O3 (2025-04-16)77.37
7Magistral Small 1.277.21
8GPT-577.16
9Yi 1.5 34B77.14
10Claude 3.7 Sonnet (20250219) (Thinking)76.61
11Qwen 3 235B A22B76.45
12GPT-5 (Minimal)76.06
13Mistral Small 376.05
14Gemini 2.5 Pro76
15Qwen 3 Next 80B A3B (Thinking)75.78

Interactive version: theaggregate.ai/benchmark?slug=euroeval-english · How It Works · Data refreshed daily, snapshot 2026-09-05.