EYT-Bench (Nemotron-USA) - Explicit Intent: leaderboard

Metric: Per-turn explicit-intent accuracy (%; the target model first predicts the user explicit-intent label from 12 classes, scored against the simulator gold label, before generating its reply; multi-turn dialogues of up to 10 turns with a persona-grounded Gemini 3.1 Pro user simulator and a Gemini 3.1 Pro judge (reasoning effort high, temperature 0); Nemotron-Personas-USA pool (100 dialogues; structured demographic personas)). Source: arxiv.org. Saturation forecast: Around April 2028. 17 models tracked.

Top models

#ModelScore
1DeepSeek V4 Pro55.2
2GPT-5.549.3
3Gemini 3.1 Pro (Preview) (Thinking)46.7
4DeepSeek V4 Flash37.4
5Qwen 3.5 397B A17B37.2
6Qwen 3.5 27B35.9
7Claude Sonnet 4.633.8
8Claude Opus 4.729.3
9Qwen 3.5 35B A3B28.3
10Seed 2.0 Mini26.4
11Seed 2.0 Pro25.2
12Seed 2.0 Lite19.6

Interactive version: theaggregate.ai/benchmark?slug=eyt-bench-nemotron-usa-explicit-intent · How It Works · Data refreshed daily, snapshot 2026-09-29.