ChartSync - Visuo-Logical Consistency: leaderboard

Metric: Visuo-logical consistency score (%; Gemini-3.1-Pro judge score in 0, 0.25, 0.5 or 1 of whether the chart geometry coupled to the edited values was updated to match, on the 235 cascading-edit instances; 870 chart editing triplets (original chart, instruction, programmatically rendered ground truth) over 9 chart categories, 235 of them geometry-coupled cascading edits; each model edits the chart image once). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 14 models tracked.

Top models

#ModelScore
1Nano Banana Pro (Gemini 3 Pro Image)83.71
2GPT-Image-274.47
3Wan2.7-Image-Pro56.38
4SeeDream-5.0-Lite37.66
5Qwen-Image-2.0-Pro24.26
6Qwen-Image-Edit-251113.83
7LongCat-Image-Edit13.19
8FireRed-Image-Edit-1.112.77
9FLUX.2-klein-base-9B12.55
10Intern-VL-U12.02

Interactive version: theaggregate.ai/benchmark?slug=chartsync-visuo-logical-consistency · How It Works · Data refreshed daily, snapshot 2026-09-29.