MechVQA - Item Localization: leaderboard

Metric: Accuracy (%) on the Item Localization subtask (Recognition capability), MechVQA test split (drawing-level 8:1:1 split of 20,778 questions on 3,281 mechanical drawings), answers judged against the reference by three LLM judges (GPT-OSS-120B, DeepSeek-V3.2, Kimi-k2); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 15 models tracked.

Top models

#ModelScore
1Gemini 3 Pro (Preview)64.03
2GLM-4.6V63.31
3GPT-562.59
4Qwen 3 VL 32B Instruct56.83
5Claude Sonnet 4.556.12
6GPT-4o40.29
7Qwen 3 VL 30B A3B Instruct37.41
8Qwen 3 VL 4B Instruct33.09
9Gemma 3 27B (IT)24.46
10GPT-4o Mini23.02
11Llama 3.2 11B Instruct11.51

Interactive version: theaggregate.ai/benchmark?slug=mechvqa-item-localization · How It Works · Data refreshed daily, snapshot 2026-10-07.