WeGenBench - Text Rendering: leaderboard

Metric: Character-level F1 (0-1 scaled to %) between the text rendered in the image (read by the TextPecker VLM reader) and the target text; images generated from the 2,000 bilingual WeGenBench-Text prompts that require rendering given text; API refusals are left out of the average; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 18 models tracked.

Top models

#ModelScore
1Seedream-4.579
2GLM-Image78
3ERNIE-Image-Turbo (No Prompt Enhancer)75
4SenseNova-U1-8B-MoT74
5Z-Image-Turbo73
6SenseNova-U1-8B-MoT (Thinking)72
7HunyuanImage-3.0-Instruct-Distil (Image)70
8HunyuanImage-3.0-Instruct (Image)69
9ERNIE-Image-Turbo (Prompt Enhancer)65
10HunyuanImage-3.0-Instruct (Think-Rewrite)63

Interactive version: theaggregate.ai/benchmark?slug=wegenbench-text-rendering · How It Works · Data refreshed daily, snapshot 2026-09-29.