VectorGym - SVG Editing: leaderboard

Metric: SVG Editing task score (0-100): applying a human-written multi-step edit to an SVG, mean of the VLM-judge score, DINO similarity, 100 minus MSE and 100 minus LPIPS of the rendered output against the human-made target (all 0-100), on VectorGym's human-annotated test split (real-world SVGs from SVG-Stack); GPT-5.1 judges; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 10 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3 Pro88.71#77
2Claude Sonnet 4.588.07#138
3GPT-5.187.71#131
4GPT-4o82.35#333
5Gemini 2.5 Flash81.3#237
6Qwen 3 VL 235B A22B Instruct80.11#264
7Qwen 3 VL 8B Instruct77.89#401
8GLM-4.5V68.34#339
9Qwen 2.5 VL 32B Instruct59.61#443
10Qwen 2.5 VL 72B Instruct57.52#364

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=vectorgym-svg-editing · How It Works · Data refreshed daily, snapshot 2026-10-11.