EVALITA - hate-speech-detection — leaderboard

Metric: CPS. Source: huggingface.co. 48 models tracked.

Top models

#ModelScore
1Llama 3.3 70B Instruct77.82
2Mistral Large 2 (Nov) Instruct (2411)77.03
3Qwen 2.5 72B Instruct76.38
4Gemma 3 27B (IT)75.87
5calme-3.2-instruct-78B75.82
6Qwen 2.5 VL 32B Instruct75.47
7DeepSeek R1 Distill Llama 70B74.6
8Gemma 2 27B (IT)74.14
9Mistral Small 372.9
10Qwen2.5-14B-Instruct-1M72.8
11Gemma 2 9B (IT)71.99
12Gemma 3 12B (IT)71.4
13LLaMAntino-3-ANITA-8B-Inst-DPO-ITA70.99
14Qwen 3 Next 80B A3B Instruct69.84
15Llama-3.1-SuperNova-Lite69.24

Interactive version: theaggregate.ai/benchmark?slug=evalita-hate-speech-detection · How the rankings work · Data refreshed daily, snapshot 2026-07-22.