TextSculpt-Bench - Text Accuracy: leaderboard

Metric: Text accuracy (0-100): one minus the full-image word edit distance between the text GPT-5.2 infers should appear after editing and the text actually observed, normalized by the edited words and clipped at 0, over 800 TextSculpt-Bench text-editing tasks (200 each of addition, removal, replacement and hybrid editing) on Pexels photos and generated posters, default model settings; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 9 models tracked.

Top models

#ModelScore
1Seedream 4.576
2Gemini-2.5-Flash-Image63
3FireRed-Image-Edit-1.062
4Seedream 4.056
5Qwen-Image-Edit-251155
6LongCat-Image-Edit54
7Step1X-Edit47
8Bagel26
9OmniGen221

Interactive version: theaggregate.ai/benchmark?slug=textsculpt-bench-text-accuracy · How It Works · Data refreshed daily, snapshot 2026-10-07.