ROSE - Global Count: leaderboard

Metric: Strict PASS (%) on T1 global counting: the number of exception cells in the whole grid, macro-averaged over the five visual sources, on the 5,560-instance test split of ROSE v0.1: grid images of repeated majority elements with a few exceptions, from five controlled visual sources (confusable Chinese glyphs, emoji styles and identities, pixel-art edits and assets); strict PASS requires an exact count or an exact set of clicked grid coordinates in the prescribed answer grammar; temperature 0, reasoning or thinking disabled where supported; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 9 models tracked.

Top models

#ModelScore
1GPT-5.5 (Non-reasoning)93.8
2Gemini 3.1 Pro (Preview)92.8
3Qwen 3.6 Plus (Non-reasoning)80.3
4Claude Opus 4.864
5Claude Sonnet 4.662.1
6GLM-4.6V (Non-reasoning)60.7

Interactive version: theaggregate.ai/benchmark?slug=rose-global-count · How It Works · Data refreshed daily, snapshot 2026-09-29.