DrawingVQA - Contextual Interpretation: leaderboard

Metric: Accuracy (%; reasoning depth R2, 34 contextual-interpretation questions of the 92 expert-curated questions on 33 Issued for Construction drawings, multiple choice and open-ended). Source: arxiv.org. Saturation forecast: Around December 2026. 19 models tracked.

Top models

#ModelScore
1Gemini 3 Pro (Preview)79.4
2Qwen 3 VL 32B Instruct79.4
3Gemini 2.5 Pro76.5
4Gemini 2.5 Flash73.5
5Claude Haiku 4.573.5
6Claude Sonnet 4.564.7
7O3 (2025-04-16)61.8
8Qwen 3 VL 8B Instruct61.8
9Llama 4 Scout Instruct58.8
10InternVL3.5-8B58.8
11Qwen 3 VL 30B A3B Instruct58.8
12Llama 3.2 11B Instruct47.1
13Phi-4 Multimodal Instruct47.1
14GPT-4o (2024-08-06)44.1

Interactive version: theaggregate.ai/benchmark?slug=drawingvqa-contextual-interpretation · How It Works · Data refreshed daily, snapshot 2026-09-29.