CNSL-bench - Illustrative Images: leaderboard
Metric: Accuracy (%) identifying the sign's meaning from the dictionary's illustrative image of the sign, four-way multiple choice over the 6,707 sign entries of the National Common Sign Language Dictionary (distractors sampled at random), official sampling settings per model, accuracy by exact match of the chosen option; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 23 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5 | 66.96 |
| 2 | Gemini 2.5 Pro (Medium) | 61.13 |
| 3 | Gemini 2.5 Flash | 51.62 |
| 4 | Gemini 2.5 Flash (Non-reasoning) | 43.57 |
| 5 | GLM-4.1V-9B (Thinking) | 39.62 |
| 6 | GPT-4o | 39.07 |
| 7 | Qwen 3 VL 8B Instruct | 38.39 |
| 8 | Qwen 3 VL 8B (Thinking) | 37.56 |
| 9 | GPT-4o Mini | 35.99 |
| 10 | Qwen 2.5 VL 7B Instruct | 33.32 |
| 11 | Qwen 2 VL 7B Instruct | 32.44 |
Interactive version: theaggregate.ai/benchmark?slug=cnsl-bench-illustrative-images · How It Works · Data refreshed daily, snapshot 2026-10-07.