StructEval — leaderboard

Structured-output benchmark evaluating text and visual structured generation and conversion across 18 formats and 2,035 examples.

Metric: Average (self-reported). Source: benchmarklist.com. Status: saturation imminent. 12 models tracked.

Top models

#ModelScore
1GPT-4o76.02
2GPT-4.1 Mini75.64
3O1 Mini75.58
4GPT-4o Mini73.19
5Gemini 1.5 Pro71.75
6Qwen 3 4B67.04
7Gemini 2.0 Flash62.55
8Llama 3.1 8B Instruct61.77
9Qwen 2.5 7B Instruct59.03
10Phi-4 Mini Instruct56.97
11Llama 3 8B Instruct51.59
12Phi-3 Mini 128K Instruct40.79

Interactive version: theaggregate.ai/benchmark?slug=structeval · How the rankings work · Data refreshed daily, snapshot 2026-07-22.