EuroEval Lithuanian NLU - WikiANN LT — leaderboard

Metric: Named entity recognition Score (%). Source: euroeval.com. 230 models tracked.

Top models

#ModelScore
1Llama 3.1 70B70.1
2Claude Sonnet 4.5 (Thinking)69.92
3gemma-3-27B-pt69.51
4Claude Sonnet 4.5 (Non-reasoning)68.55
5GPT-5.4 Mini (High)68.36
6GPT-5.4 Mini (Medium)67.66
7Mistral Small 3.267.61
8Gemma 4 31B (IT)67.18
9GPT-567.12
10Mistral Small 3.167.09
11Qwen 3 Next 80B A3B Instruct66.42
12Llama 3.3 70B Instruct66.31
13Ministral 3 14B66.18
14Qwen 3 235B A22B 2507 Instruct66.15
15Qwen 3 Next 80B A3B (Thinking)65.98

Interactive version: theaggregate.ai/benchmark?slug=euroeval-lithuanian-nlu-wikiann-lt · How the rankings work · Data refreshed daily, snapshot 2026-07-22.