OpenHuEval — leaderboard
OpenHuEval evaluates large language models on Hungarian-specific tasks, including real user queries, self-awareness, proverb reasoning, generative evaluation, and fill-in-the-blank tasks.
Metric: Macro Average (computed) (self-reported). Source: benchmarklist.com. Status: saturation imminent. 10 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-4o | 63.77 |
| 2 | DeepSeek V3 | 57.1 |
| 3 | Llama 3.1 70B Instruct | 50.41 |
| 4 | O1 Mini | 49.38 |
| 5 | GPT-4o Mini | 49.33 |
| 6 | Llama 3.1 8B Instruct | 28.73 |
Interactive version: theaggregate.ai/benchmark?slug=openhueval · How the rankings work · Data refreshed daily, snapshot 2026-07-22.