ROSE - Chinese Glyphs: leaderboard

Metric: Strict PASS (%) averaged over the five coupled tasks on the ChineseGlyph source (confusable characters in one verified font), on the 5,560-instance test split of ROSE v0.1: grid images of repeated majority elements with a few exceptions, from five controlled visual sources (confusable Chinese glyphs, emoji styles and identities, pixel-art edits and assets); strict PASS requires an exact count or an exact set of clicked grid coordinates in the prescribed answer grammar; temperature 0, reasoning or thinking disabled where supported; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 9 models tracked.

Top models

#ModelScore
1GPT-5.5 (Non-reasoning)87.4
2Gemini 3.1 Pro (Preview)67.6
3Qwen 3.6 Plus (Non-reasoning)48.4
4Claude Opus 4.830.2
5Claude Sonnet 4.628.5
6GLM-4.6V (Non-reasoning)19.1

Interactive version: theaggregate.ai/benchmark?slug=rose-chinese-glyphs · How It Works · Data refreshed daily, snapshot 2026-09-29.