AesCanvas - Contextual Suitability: leaderboard

Metric: Exact-match accuracy (%; one option label per ContextCanvas case on whether an image is an appropriate visual choice for a stated use scenario, audience and context; all 301 expert-reviewed cases (291 two-option, 10 three-option); images resized to 448 x 448, deterministic decoding, no retrieval or tools, unparsable answers counted wrong). Source: arxiv.org. Saturation forecast: Around December 2026. 17 models tracked.

Top models

#ModelScore
1Claude Opus 591.69
2Gemini 3.1 Pro (Preview)85.71
3GPT-5.583.06
4Claude Opus 4.672.76
5GPT-5.271.43
6GLM-5V Turbo57.81
7Qwen 3.7 Plus51.83
8Grok 4.2044.19
9InternVL3-8B23.26

Interactive version: theaggregate.ai/benchmark?slug=aescanvas-contextual-suitability · How It Works · Data refreshed daily, snapshot 2026-09-26.