Arena-T2I Hard: leaderboard

Metric: Faithfulness: dependency-aware faithfulness yes-ratio (%; share of the about 30 decomposed yes/no constraint questions per prompt that a gemini-3-flash judge answers yes for the generated image, a question counted only when its parent questions pass) on the 310 Arena-T2I Hard prompts drawn from real arena text-to-image logs; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 13 models tracked.

Top models

#ModelScore
1gemini-3-pro-image-preview-2k85.5
2Grok Imagine Image 2026030684.9
3gpt-image-1.5-high-fidelity79.6
4recraft-v478.7
5wan2.6-t2i-v276.8
6gemini-2.5-flash-image76.8
7gpt-image-172.2
8imagen-4.0-ultra-generate-00168
9imagen-4.0-generate-00165.9
10hunyuan-image-3.0-fal60.9

Interactive version: theaggregate.ai/benchmark?slug=arena-t2i-hard · How It Works · Data refreshed daily, snapshot 2026-09-29.