Arena-T2I Hard: leaderboard
Metric: Faithfulness: dependency-aware faithfulness yes-ratio (%; share of the about 30 decomposed yes/no constraint questions per prompt that a gemini-3-flash judge answers yes for the generated image, a question counted only when its parent questions pass) on the 310 Arena-T2I Hard prompts drawn from real arena text-to-image logs; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 13 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | gemini-3-pro-image-preview-2k | 85.5 |
| 2 | Grok Imagine Image 20260306 | 84.9 |
| 3 | gpt-image-1.5-high-fidelity | 79.6 |
| 4 | recraft-v4 | 78.7 |
| 5 | wan2.6-t2i-v2 | 76.8 |
| 6 | gemini-2.5-flash-image | 76.8 |
| 7 | gpt-image-1 | 72.2 |
| 8 | imagen-4.0-ultra-generate-001 | 68 |
| 9 | imagen-4.0-generate-001 | 65.9 |
| 10 | hunyuan-image-3.0-fal | 60.9 |
Interactive version: theaggregate.ai/benchmark?slug=arena-t2i-hard · How It Works · Data refreshed daily, snapshot 2026-09-29.