ChartSync - Visuo-Logical Consistency: leaderboard
Metric: Visuo-logical consistency score (%; Gemini-3.1-Pro judge score in 0, 0.25, 0.5 or 1 of whether the chart geometry coupled to the edited values was updated to match, on the 235 cascading-edit instances; 870 chart editing triplets (original chart, instruction, programmatically rendered ground truth) over 9 chart categories, 235 of them geometry-coupled cascading edits; each model edits the chart image once). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 14 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Nano Banana Pro (Gemini 3 Pro Image) | 83.71 |
| 2 | GPT-Image-2 | 74.47 |
| 3 | Wan2.7-Image-Pro | 56.38 |
| 4 | SeeDream-5.0-Lite | 37.66 |
| 5 | Qwen-Image-2.0-Pro | 24.26 |
| 6 | Qwen-Image-Edit-2511 | 13.83 |
| 7 | LongCat-Image-Edit | 13.19 |
| 8 | FireRed-Image-Edit-1.1 | 12.77 |
| 9 | FLUX.2-klein-base-9B | 12.55 |
| 10 | Intern-VL-U | 12.02 |
Interactive version: theaggregate.ai/benchmark?slug=chartsync-visuo-logical-consistency · How It Works · Data refreshed daily, snapshot 2026-09-29.