ROSE - Emoji: leaderboard

Metric: Strict PASS (%) averaged over the five coupled tasks on the two emoji sources (same emoji in different renderings; related emoji in one style), averaged, on the 5,560-instance test split of ROSE v0.1: grid images of repeated majority elements with a few exceptions, from five controlled visual sources (confusable Chinese glyphs, emoji styles and identities, pixel-art edits and assets); strict PASS requires an exact count or an exact set of clicked grid coordinates in the prescribed answer grammar; temperature 0, reasoning or thinking disabled where supported; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 9 models tracked.

Top models

#ModelScore
1GPT-5.5 (Non-reasoning)94.8
2Gemini 3.1 Pro (Preview)84.5
3Qwen 3.6 Plus (Non-reasoning)48
4Claude Opus 4.825.2
5Claude Sonnet 4.625
6GLM-4.6V (Non-reasoning)22.4

Interactive version: theaggregate.ai/benchmark?slug=rose-emoji · How It Works · Data refreshed daily, snapshot 2026-09-29.