EYT-Bench (Nemotron-USA) - Latent Intent: leaderboard

Metric: Per-turn latent-intent accuracy (%; the target model first predicts the user latent-intent label from 8 classes, scored against the simulator gold label, before generating its reply; multi-turn dialogues of up to 10 turns with a persona-grounded Gemini 3.1 Pro user simulator and a Gemini 3.1 Pro judge (reasoning effort high, temperature 0); Nemotron-Personas-USA pool (100 dialogues; structured demographic personas)). Source: arxiv.org. Saturation forecast: Around July 2027. 17 models tracked.

Top models

#ModelScore
1DeepSeek V4 Pro57.6
2GPT-5.549.2
3Gemini 3.1 Pro (Preview) (Thinking)46.1
4DeepSeek V4 Flash39.8
5Qwen 3.5 397B A17B36.3
6Claude Opus 4.734.2
7Qwen 3.5 27B28.5
8Claude Sonnet 4.627
9Qwen 3.5 35B A3B25.7
10Seed 2.0 Pro22.5
11Seed 2.0 Lite20.9
12Seed 2.0 Mini16

Interactive version: theaggregate.ai/benchmark?slug=eyt-bench-nemotron-usa-latent-intent · How It Works · Data refreshed daily, snapshot 2026-09-29.