ROSE - Pixel Art: leaderboard

Metric: Strict PASS (%) averaged over the five coupled tasks on the two pixel-art sources (localized edits of one asset; related assets), averaged, on the 5,560-instance test split of ROSE v0.1: grid images of repeated majority elements with a few exceptions, from five controlled visual sources (confusable Chinese glyphs, emoji styles and identities, pixel-art edits and assets); strict PASS requires an exact count or an exact set of clicked grid coordinates in the prescribed answer grammar; temperature 0, reasoning or thinking disabled where supported; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 9 models tracked.

Top models

#ModelScore
1GPT-5.5 (Non-reasoning)92.2
2Gemini 3.1 Pro (Preview)80.1
3Qwen 3.6 Plus (Non-reasoning)53.6
4Claude Opus 4.820.4
5Claude Sonnet 4.620.1
6GLM-4.6V (Non-reasoning)19.9

Interactive version: theaggregate.ai/benchmark?slug=rose-pixel-art · How It Works · Data refreshed daily, snapshot 2026-09-29.