EuroEval Spanish NLU — leaderboard

Metric: NLU Average Score (%). Source: euroeval.com. 418 models tracked.

Top models

#ModelScore
1Gemini 3 Pro (Preview)65.42
2Claude Sonnet 4.665.08
3GPT-5.4 Mini (High)63.97
4GPT-5.4 Mini (Medium)62.2
5Gemini 3 Flash (Preview) (Non-reasoning)62.13
6GPT-5.4 Mini (Non-reasoning)61.33
7Qwen 3 4B 2507 (Thinking)60.1
8Grok 4.1 Fast (Reasoning)60.07
9Gemini 3.1 Flash Lite (Preview)58.92
10GPT-5.258.63
11Mistral Small 3.258.35
12Mistral Small 3.158.31
13GPT-5.4 Nano (Medium)58.08
14Qwen 3 Next 80B A3B Instruct58.04
15Qwen 3 235B A22B 2507 Instruct57.84

Interactive version: theaggregate.ai/benchmark?slug=euroeval-spanish-nlu · How the rankings work · Data refreshed daily, snapshot 2026-07-22.