E-comIQ-18k - Text: leaderboard

Metric: Acc@0.5 (%): share of the 1,000 E-comIQ-18k test posters (AI-generated Chinese e-commerce posters) whose predicted text-rendering score falls within 0.5 of the expert-calibrated reference score on the 1-5 quality scale; higher is better. Source: arxiv.org. Saturation forecast: Around 2034. 11 models tracked.

Top models

#ModelScoreOverall rank
1GPT-4o34.3#333
2Qwen 2.5 VL 7B Instruct30.1#643
3Gemini 2.5 Pro29.2#145
4Claude Sonnet 4.527.8#138
5Grok 426.9#169
6Qwen 2.5 VL 72B Instruct25.2#364

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=e-comiq-18k-text · How It Works · Data refreshed daily, snapshot 2026-10-11.