Hy-MultiTurn - Constraint Synthesis: leaderboard

Metric: Importance-weighted score (%) on the 35 constraint synthesis dialogues, 209 long Chinese multi-turn dialogues (12 to 76 turns), three rolls per task averaged, three LLM judges (GPT-5.2, Qwen 3.6, DeepSeek-V4 Pro) averaged, a deterministic checker for precise execution; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 22 models tracked.

Top models

#ModelScore
1Grok 4.5 (High)90
2GPT-5.5 (High)84.9
3GPT-5.4 (High)83.7
4Claude Opus 4.6 (High)83.4
5Kimi K3 (Max)80.2
6Hy3 (High)79.5
7Claude Opus 4.7 (High)74.9
8Kimi K2.668.8
9Doubao-Seed-2.1-Pro (High)65.3
10Hy3-preview (High)62.6
11Kimi K2.6 (Non-reasoning)61.5
12Seed 2.0 Pro (Thinking)59.7
13Kimi K2.5 (Non-reasoning)59.3
14Qwen 3.6 Plus (Thinking)56.9
15Hy3-preview (Non-reasoning)55.6

Interactive version: theaggregate.ai/benchmark?slug=hy-multiturn-constraint-synthesis · How It Works · Data refreshed daily, snapshot 2026-09-29.