EYT-Bench (Nemotron-USA) - Anthropomorphic Interaction: leaderboard

Metric: Anthropomorphic interaction (colloquial, emotional, flow, flexibility, rhythm) (0-100): judge rubric of five sub-indicators scored 0, 0.5 or 1 per turn, summed to 0-5, turn weighted with the first two turns at 0.10 each and rescaled to 0-100; multi-turn dialogues of up to 10 turns with a persona-grounded Gemini 3.1 Pro user simulator and a Gemini 3.1 Pro judge (reasoning effort high, temperature 0); Nemotron-Personas-USA pool (100 dialogues; structured demographic personas). Source: arxiv.org. Saturation forecast: Estimated already saturated. 17 models tracked.

Top models

#ModelScore
1Seed 2.0 Pro96
2Claude Sonnet 4.695.9
3Claude Opus 4.794.9
4Seed 2.0 Lite94.2
5DeepSeek V4 Pro92.9
6Qwen 3.5 27B92.4
7Gemini 3.1 Pro (Preview) (Thinking)90.9
8Qwen 3.5 397B A17B90.7
9DeepSeek V4 Flash89.2
10Qwen 3.5 35B A3B88.2
11GPT-5.575.6
12Seed 2.0 Mini75.4

Interactive version: theaggregate.ai/benchmark?slug=eyt-bench-nemotron-usa-anthropomorphic-interaction · How It Works · Data refreshed daily, snapshot 2026-09-29.