WeGenBench - Text Rendering Sentence Accuracy: leaderboard

Metric: Sentence accuracy (0-1 scaled to %), the exact-match rate of rendered text segments against the target text; images generated from the 2,000 bilingual WeGenBench-Text prompts that require rendering given text; API refusals are left out of the average; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 18 models tracked.

Top models

#ModelScore
1Seedream-4.559
2GPT-Image-257
3GLM-Image56
4HunyuanImage-3.0-Instruct (Image)56
5HunyuanImage-3.0-Instruct (Think-Rewrite)56
6ERNIE-Image-Turbo (Prompt Enhancer)56
7ERNIE-Image-Turbo (No Prompt Enhancer)56
8Nano-banana-255
9HunyuanImage-3.0-Instruct-Distil (Think-Rewrite)54
10HiDream-O1-Image54

Interactive version: theaggregate.ai/benchmark?slug=wegenbench-text-rendering-sentence-accuracy · How It Works · Data refreshed daily, snapshot 2026-09-29.