Polish EQ-Bench — leaderboard

Polish adaptation of EQ-Bench: emotional intelligence benchmark testing LLMs on understanding and predicting emotional dynamics in Polish-language scenarios.

Metric: EQ-Bench Score. Source: huggingface.co. Status: saturated. 102 models tracked.

Top models

#ModelScore
1Mistral Large 2 (Jul)78.07
2GPT-4 Turbo77.77
3Mistral Large 2 (Nov) Instruct (2411)77.29
4Llama 3.1 405B Instruct FP877.23
5GPT-4o (2024-08-06)75.15
6DeepSeek V3 (0324)73.46
7Llama 3.3 70B Instruct72.86
8Mistral-Small-Instruct-240972.85
9Llama 3.1 70B Instruct72.53
10Qwen 2 72B Instruct72.07
11Llama 3 70B Instruct71.21
12Qwen 2.5 32B Instruct71.15
13GPT-4o Mini (2024-07-18)71.15
14Bielik-11B-v2.3-Instruct70.86
15Mistral Small 370.52

Interactive version: theaggregate.ai/benchmark?slug=polish-eq-bench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.