EVALITA - word-in-context — leaderboard

Metric: CPS. Source: huggingface.co. 48 models tracked.

Top models

#ModelScore
1Qwen 2.5 VL 32B Instruct77.56
2Mistral Large 2 (Nov) Instruct (2411)75.55
3calme-3.2-instruct-78B73.44
4Qwen 2.5 72B Instruct72.06
5Mistral Small 371.65
6Qwen 3 Next 80B A3B Instruct71.41
7DeepSeek R1 Distill Llama 70B69.26
8Phi-3-medium-4k-instruct68.92
9Gemma 2 27B (IT)68.56
10Llama 3.3 70B Instruct67.15
11Gemma 3 12B (IT)66.5
12Gemma 3 27B (IT)66.49
13Llama 3 8B 4bit UltraChat Ita66.22
14Phi-3.5-mini-instruct65.33
15granite-3.1-8B-instruct65.03

Interactive version: theaggregate.ai/benchmark?slug=evalita-word-in-context · How the rankings work · Data refreshed daily, snapshot 2026-07-22.