TextSculpt-Bench - Visual Quality: leaderboard

Metric: Visual quality (0-100): mean of three yes/no GPT-5.2 judgments (edit location, style consistency, physical plausibility), over 800 TextSculpt-Bench text-editing tasks (200 each of addition, removal, replacement and hybrid editing) on Pexels photos and generated posters, default model settings; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 9 models tracked.

Top models

#ModelScore
1Seedream 4.576
2FireRed-Image-Edit-1.070
3Gemini-2.5-Flash-Image69
4Qwen-Image-Edit-251163
5Seedream 4.057
6LongCat-Image-Edit54
7Step1X-Edit45
8OmniGen228
9Bagel25

Interactive version: theaggregate.ai/benchmark?slug=textsculpt-bench-visual-quality · How It Works · Data refreshed daily, snapshot 2026-10-07.