StickToYourRole — leaderboard

Role-playing consistency benchmark: measures how stably LLMs maintain simulated personal values during multi-turn role-playing conversations using Schwartz value theory.

Metric: Cardinal Score. Source: huggingface.co. Status: saturation imminent. 32 models tracked.

Top models

#ModelScore
1Llama 3.1 Nemotron 70B Instruct80.66
2Mistral Large 2 (Jul)78.65
3Llama 3.3 70B Instruct78.26
4QwQ-32B77.19
5Llama 3.1 70B Instruct77.16
6Qwen 3 32B73.64
7Mistral Large 2 (Nov) Instruct (2411)73.35
8Qwen 3 8B71.84
9Qwen 3 235B A22B FP871.84
10Mistral Small 3.170.26
11Qwen2.5-14B-Instruct-1M70.22
12Qwen 3 4B69.71
13Llama 3.1 8B Instruct62
14Llama 4 Scout Instruct61.81
15Cydonia-22B-v1.256.6

Interactive version: theaggregate.ai/benchmark?slug=sticktoyourrole · How the rankings work · Data refreshed daily, snapshot 2026-07-22.