AlpsBench - Persona Awareness: leaderboard

Metric: Pass rate (0-1, times 100) of responses to user queries that must recall and apply explicit user attributes, judged pass or fail by a DeepSeek-V3.2 judge given the dialogue history; AlpsBench's 2,500 human-verified instances per task built from long-term real user-LLM dialogues in WildChat; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 7 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3 Flash (Preview)68.89#78
2Qwen 3 Max62.46#201
3DeepSeek V3.2 (Thinking)59.12#198 (DeepSeek V3.2)
4GPT-5.256.84#105
5Claude Sonnet 4.556.49#138
6GPT-4.1 Mini41.12#346
7Llama 4 Maverick27.89#451

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=alpsbench-persona-awareness · How It Works · Data refreshed daily, snapshot 2026-10-11.