StickToYourRole — leaderboard
Role-playing consistency benchmark: measures how stably LLMs maintain simulated personal values during multi-turn role-playing conversations using Schwartz value theory.
Metric: Cardinal Score. Source: huggingface.co. Status: saturation imminent. 32 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Llama 3.1 Nemotron 70B Instruct | 80.66 |
| 2 | Mistral Large 2 (Jul) | 78.65 |
| 3 | Llama 3.3 70B Instruct | 78.26 |
| 4 | QwQ-32B | 77.19 |
| 5 | Llama 3.1 70B Instruct | 77.16 |
| 6 | Qwen 3 32B | 73.64 |
| 7 | Mistral Large 2 (Nov) Instruct (2411) | 73.35 |
| 8 | Qwen 3 8B | 71.84 |
| 9 | Qwen 3 235B A22B FP8 | 71.84 |
| 10 | Mistral Small 3.1 | 70.26 |
| 11 | Qwen2.5-14B-Instruct-1M | 70.22 |
| 12 | Qwen 3 4B | 69.71 |
| 13 | Llama 3.1 8B Instruct | 62 |
| 14 | Llama 4 Scout Instruct | 61.81 |
| 15 | Cydonia-22B-v1.2 | 56.6 |
Interactive version: theaggregate.ai/benchmark?slug=sticktoyourrole · How the rankings work · Data refreshed daily, snapshot 2026-07-22.