ChartSync: leaderboard

Metric: Overall score (%; equal average of five metric-level scores: character-level OCR F1 of all chart text, SSIM against the ground truth, and three Gemini-3.1-Pro judge scores for textual edit success, visuo-logical consistency and background fidelity; 870 chart editing triplets (original chart, instruction, programmatically rendered ground truth) over 9 chart categories, 235 of them geometry-coupled cascading edits; each model edits the chart image once). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 14 models tracked.

Top models

#ModelScore
1Nano Banana Pro (Gemini 3 Pro Image)87.76
2GPT-Image-281.78
3Wan2.7-Image-Pro74.47
4SeeDream-5.0-Lite71.09
5Qwen-Image-2.0-Pro65.54
6Qwen-Image-Edit-251160.29
7FireRed-Image-Edit-1.156.22
8FireRed-Image-Edit-1.051.61
9Qwen-Image-Edit50.75
10FLUX.2-klein-base-9B50.62

Interactive version: theaggregate.ai/benchmark?slug=chartsync · How It Works · Data refreshed daily, snapshot 2026-09-29.