MechVQA - Text and Table: leaderboard

Metric: Accuracy (%) on the Text and Table subtask (Recognition capability), MechVQA test split (drawing-level 8:1:1 split of 20,778 questions on 3,281 mechanical drawings), answers judged against the reference by three LLM judges (GPT-OSS-120B, DeepSeek-V3.2, Kimi-k2); higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 15 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.597.73
2Gemini 3 Pro (Preview)97.73
3Qwen 3 VL 32B Instruct97.73
4GLM-4.6V97.73
5GPT-593.18
6Qwen 3 VL 30B A3B Instruct90.91
7Qwen 3 VL 4B Instruct88.64
8GPT-4o81.82
9Gemma 3 27B (IT)70.45
10GPT-4o Mini68.18
11Llama 3.2 11B Instruct54.55

Interactive version: theaggregate.ai/benchmark?slug=mechvqa-text-and-table · How It Works · Data refreshed daily, snapshot 2026-10-07.