LaViSA - Per-Sentence Accuracy: leaderboard

Metric: Per-sentence accuracy (%): share of the 700 ambiguous sentences for which the model picks the correct reading for every one of its images, averaged over the seven ambiguity categories; task: pick which disambiguated reading of a structurally ambiguous sentence a generated image depicts (2 or 3 readings per sentence, 100 sentences per category); higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 13 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)77.7
2Gemini 3.1 Flash Lite69
3GPT-5.267.1
4Qwen 3 VL 32B Instruct64.3
5Qwen 3 VL 32B (Thinking)61.1
6Qwen 3 VL 4B Instruct57
7Qwen 3 VL 8B (Thinking)50.3
8Qwen 3 VL 8B Instruct49.4
9Qwen 3 VL 4B (Thinking)44
10Gemma 3 27B (IT)43.9
11Gemma 3 12B (IT)42.3
12Gemma 3 4B (IT)9.1

Interactive version: theaggregate.ai/benchmark?slug=lavisa-per-sentence-accuracy · How It Works · Data refreshed daily, snapshot 2026-09-29.