EVALITA - sentiment-analysis — leaderboard

Metric: CPS. Source: huggingface.co. 48 models tracked.

Top models

#ModelScore
1Mistral Large 2 (Nov) Instruct (2411)80.94
2Gemma 3 27B (IT)80.3
3calme-3.2-instruct-78B80.15
4Llama 3.3 70B Instruct79.55
5Gemma 3 12B (IT)78.34
6Qwen 2.5 72B Instruct77.68
7DeepSeek R1 Distill Llama 70B77.48
8Gemma 3n E4B (IT)76.34
9Mistral Small 375.95
10Gemma 3 4B (IT)75.3
11Llama-3.1-SuperNova-Lite75.24
12Phi-475.21
13granite-3.1-8B-instruct74.76
14Gemma 2 27B (IT)74.26
15Qwen 2.5 7B Instruct74.16

Interactive version: theaggregate.ai/benchmark?slug=evalita-sentiment-analysis · How the rankings work · Data refreshed daily, snapshot 2026-07-22.