MechVQA - Assembly Relationship: leaderboard

Metric: Accuracy (%) on the Assembly Relationship subtask (Reasoning capability), MechVQA test split (drawing-level 8:1:1 split of 20,778 questions on 3,281 mechanical drawings), answers judged against the reference by three LLM judges (GPT-OSS-120B, DeepSeek-V3.2, Kimi-k2); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 15 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.562
2Qwen 3 VL 32B Instruct62
3GPT-560
4GLM-4.6V60
5GPT-4o48
6Gemini 3 Pro (Preview)46
7Qwen 3 VL 30B A3B Instruct38
8Gemma 3 27B (IT)24
9GPT-4o Mini24
10Qwen 3 VL 4B Instruct22
11Llama 3.2 11B Instruct0

Interactive version: theaggregate.ai/benchmark?slug=mechvqa-assembly-relationship · How It Works · Data refreshed daily, snapshot 2026-10-07.