EYT-Bench (PersonaMem-v2) - Final Intent Completion: leaderboard

Metric: Final-intent completion rate (%; share of dialogues in which the judge, reading the full transcript, decides the assistant helped the user reach the stated final intent; multi-turn dialogues of up to 10 turns with a persona-grounded Gemini 3.1 Pro user simulator and a Gemini 3.1 Pro judge (reasoning effort high, temperature 0); PersonaMem-v2 pool (100 dialogues; personas are free-text paragraphs distilled from long real user-assistant interactions)). Source: arxiv.org. Saturation forecast: Around December 2026. 17 models tracked.

Top models

#ModelScore
1Seed 2.0 Lite88.6
2Seed 2.0 Mini87.9
3DeepSeek V4 Flash87.5
4Seed 2.0 Pro86.2
5Qwen 3.5 35B A3B83.3
6Qwen 3.5 397B A17B81.2
7Qwen 3.5 27B80.9
8Gemini 3.1 Pro (Preview) (Thinking)80
9DeepSeek V4 Pro78.1
10Claude Opus 4.769
11Claude Sonnet 4.666
12GPT-5.552.5

Interactive version: theaggregate.ai/benchmark?slug=eyt-bench-personamem-v2-final-intent-completion · How It Works · Data refreshed daily, snapshot 2026-09-29.