TECCI - Instruction Rich Challenge Set: leaderboard

Metric: Overall human success rate (%; five trained raters score each edit 1-5 and a criterion succeeds when the mean is at least 4.5 on all three criteria; 265 sampled pairs with manually written challenging instructions). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 5 models tracked.

Top models

#ModelScore
1Nano Banana Pro19.7
2Nano Banana 2 (Gemini 3.1 Flash Image Preview)17.7
3Grok Imagine Pro14.7
4Seedream 5.0 Lite10.3
5GPT Image 1.54.9

Interactive version: theaggregate.ai/benchmark?slug=tecci-instruction-rich-challenge-set · How It Works · Data refreshed daily, snapshot 2026-09-26.