Arena-T2I Hard - Text Rendering: leaderboard
Metric: Faithfulness on the Arena-T2I Hard prompts whose primary category is text rendering: dependency-aware faithfulness yes-ratio (%; share of the about 30 decomposed yes/no constraint questions per prompt that a gemini-3-flash judge answers yes for the generated image, a question counted only when its parent questions pass); higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 11 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | gemini-3-pro-image-preview-2k | 90.6 |
| 2 | Grok Imagine Image 20260306 | 88.1 |
| 3 | recraft-v4 | 81.7 |
| 4 | wan2.6-t2i-v2 | 80 |
| 5 | gemini-2.5-flash-image | 75 |
| 6 | gpt-image-1.5-high-fidelity | 74.4 |
| 7 | gpt-image-1 | 72.6 |
| 8 | imagen-4.0-ultra-generate-001 | 65.9 |
| 9 | imagen-4.0-generate-001 | 60.2 |
| 10 | ideogram-v3-quality | 57 |
Interactive version: theaggregate.ai/benchmark?slug=arena-t2i-hard-text-rendering · How It Works · Data refreshed daily, snapshot 2026-09-29.