EVALITA - lexical-substitution — leaderboard

Metric: CPS. Source: huggingface.co. 48 models tracked.

Top models

#ModelScore
1Qwen 3 Next 80B A3B Instruct54.53
2Mistral Small 352.67
3Qwen 2.5 72B Instruct50.24
4Llama 3.3 70B Instruct46.8
5Gemma 3 27B (IT)42.41
6Lexora-Medium-7B40.88
7Qwen 2.5 VL 32B Instruct40.67
8Llama 4 Scout Instruct39.29
9calme-3.2-instruct-78B38.82
10DeepSeek R1 Distill Llama 70B38.61
11Mistral Large 2 (Nov) Instruct (2411)38.56
12Qwen2.5-14B-Instruct-1M36.76
13Gemma 2 27B (IT)36.04
14Phi-3-medium-4k-instruct35.87
15Phi-435.1

Interactive version: theaggregate.ai/benchmark?slug=evalita-lexical-substitution · How the rankings work · Data refreshed daily, snapshot 2026-07-22.