ROSE - Emoji: leaderboard
Metric: Strict PASS (%) averaged over the five coupled tasks on the two emoji sources (same emoji in different renderings; related emoji in one style), averaged, on the 5,560-instance test split of ROSE v0.1: grid images of repeated majority elements with a few exceptions, from five controlled visual sources (confusable Chinese glyphs, emoji styles and identities, pixel-art edits and assets); strict PASS requires an exact count or an exact set of clicked grid coordinates in the prescribed answer grammar; temperature 0, reasoning or thinking disabled where supported; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.5 (Non-reasoning) | 94.8 |
| 2 | Gemini 3.1 Pro (Preview) | 84.5 |
| 3 | Qwen 3.6 Plus (Non-reasoning) | 48 |
| 4 | Claude Opus 4.8 | 25.2 |
| 5 | Claude Sonnet 4.6 | 25 |
| 6 | GLM-4.6V (Non-reasoning) | 22.4 |
Interactive version: theaggregate.ai/benchmark?slug=rose-emoji · How It Works · Data refreshed daily, snapshot 2026-09-29.