AtelierEval - Prompt Constraint Coverage - Constrained Creation (Novice Prompting): leaderboard

Metric: Prompt checklist coverage (%; share of each task's expert constraint checklist that the written prompt specifies, checked by a GPT-5.4 AtelierJudge; 120 constrained creation tasks: structured multi-constraint specifications; direct natural-language prompting). Source: arxiv.org. Saturation forecast: Around 2031. 8 models tracked.

Top models

#ModelScore
1GPT-5.237.5
2Gemini 3 Pro (Preview)36.2
3GPT-4.132.9
4Qwen 3 VL 235B A22B31.9
5GPT-4.1 Nano28.4
6Claude Sonnet 4.527.8
7Qwen 3 VL 8B26.5
8Gemini 2.0 Flash22.8

Interactive version: theaggregate.ai/benchmark?slug=ateliereval-prompt-constraint-coverage-constrained-creation-novice-prompting · How It Works · Data refreshed daily, snapshot 2026-09-26.