Structured Output Benchmark — leaderboard

Structured-output benchmark measuring schema-constrained generation with value accuracy, faithfulness, JSON validity, path recall, type safety, and perfect-output rates.

Metric: Overall (%). Source: interfaze.ai. Status: saturated. 29 models tracked.

Top models

#ModelScore
1GPT-5.487
2Gemini 3.1 Pro (Preview)86.9
3GLM-5.186.6
4Claude Opus 4.786.4
5Claude Sonnet 586.2
6GLM-4.786.1
7Qwen 3.5 35B A3B86.1
8GPT-5.586
9Gemini 2.5 Flash86
10Qwen 3 235B A22B85.7
11Claude Sonnet 4.685.4
12Claude Opus 4.685.3
13Kimi K2.685.3
14DeepSeek V4 Pro85.3
15GPT-4.185

Interactive version: theaggregate.ai/benchmark?slug=structured-output-benchmark · How the rankings work · Data refreshed daily, snapshot 2026-07-22.