RP-Bench — leaderboard

Roleplay benchmark evaluating character consistency, user agency, lorebook use, temporal reasoning, and interactive writing quality.

Metric: Combined Score. Source: huggingface.co. Status: saturation imminent. 8 models tracked.

Top models

#ModelScore
1Claude Opus 4.682
2DeepSeek V3.281.2
3GPT-4.180.2
4GLM-4.779.7
5Claude Sonnet 4.579.2
6Gemini 2.5 Flash76.4
7mistral-small-creative75.7
8qwen3.5-flash75.5

Interactive version: theaggregate.ai/benchmark?slug=rp-bench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.