Psi-Bench - Everyday Request Personalization: leaderboard

Metric: Personalized response score (how well responses are tailored to the hidden client profile), everyday requests (1-9; DeepSeek-v3.2 judge, which also plays the persona-profiled client; three rounds after the client's opening message, client profile hidden from the model). Source: arxiv.org. Saturation forecast: Around 2028. 10 models tracked.

Top models

#ModelScore
1GPT-5.16.12
2Qwen 3 Next 80B A3B5.77
3DeepSeek V4 Pro5.15
4Gemini 3.1 Pro (Preview)5.14
5Gemini 3 Flash5.1
6Qwen 3 32B4.66
7GPT-5 Mini4.42
8Grok 4 Fast4.34
9DeepSeek V3.24.32
10Qwen 3 8B3.63

Interactive version: theaggregate.ai/benchmark?slug=psi-bench-everyday-request-personalization · How It Works · Data refreshed daily, snapshot 2026-09-29.