EYT-Bench (PersonaMem-v2) - Latent Intent: leaderboard

Metric: Per-turn latent-intent accuracy (%; the target model first predicts the user latent-intent label from 8 classes, scored against the simulator gold label, before generating its reply; multi-turn dialogues of up to 10 turns with a persona-grounded Gemini 3.1 Pro user simulator and a Gemini 3.1 Pro judge (reasoning effort high, temperature 0); PersonaMem-v2 pool (100 dialogues; personas are free-text paragraphs distilled from long real user-assistant interactions)). Source: arxiv.org. Saturation forecast: Around December 2026. 17 models tracked.

Top models

#ModelScore
1DeepSeek V4 Pro79.9
2DeepSeek V4 Flash70.1
3GPT-5.547.9
4Gemini 3.1 Pro (Preview) (Thinking)47.3
5Claude Opus 4.741.5
6Claude Sonnet 4.624.9
7Qwen 3.5 397B A17B15.3
8Qwen 3.5 35B A3B14.2
9Qwen 3.5 27B13.1
10Seed 2.0 Pro12
11Seed 2.0 Mini11.7
12Seed 2.0 Lite7.7

Interactive version: theaggregate.ai/benchmark?slug=eyt-bench-personamem-v2-latent-intent · How It Works · Data refreshed daily, snapshot 2026-09-29.