AesCanvas - Contextual Suitability: leaderboard
Metric: Exact-match accuracy (%; one option label per ContextCanvas case on whether an image is an appropriate visual choice for a stated use scenario, audience and context; all 301 expert-reviewed cases (291 two-option, 10 three-option); images resized to 448 x 448, deterministic decoding, no retrieval or tools, unparsable answers counted wrong). Source: arxiv.org. Saturation forecast: Around December 2026. 17 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 5 | 91.69 |
| 2 | Gemini 3.1 Pro (Preview) | 85.71 |
| 3 | GPT-5.5 | 83.06 |
| 4 | Claude Opus 4.6 | 72.76 |
| 5 | GPT-5.2 | 71.43 |
| 6 | GLM-5V Turbo | 57.81 |
| 7 | Qwen 3.7 Plus | 51.83 |
| 8 | Grok 4.20 | 44.19 |
| 9 | InternVL3-8B | 23.26 |
Interactive version: theaggregate.ai/benchmark?slug=aescanvas-contextual-suitability · How It Works · Data refreshed daily, snapshot 2026-09-26.