MechVQA - Anomaly Detection: leaderboard

Metric: Accuracy (%) on the Anomaly Detection subtask (Judging capability), MechVQA test split (drawing-level 8:1:1 split of 20,778 questions on 3,281 mechanical drawings), answers judged against the reference by three LLM judges (GPT-OSS-120B, DeepSeek-V3.2, Kimi-k2); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 15 models tracked.

Top models

#ModelScore
1Gemini 3 Pro (Preview)78.37
2Qwen 3 VL 32B Instruct75.92
3GLM-4.6V74.29
4GPT-571.02
5Claude Sonnet 4.564.9
6Qwen 3 VL 30B A3B Instruct64.08
7Qwen 3 VL 4B Instruct62.86
8GPT-4o53.06
9GPT-4o Mini35.1
10Gemma 3 27B (IT)24.9
11Llama 3.2 11B Instruct17.55

Interactive version: theaggregate.ai/benchmark?slug=mechvqa-anomaly-detection · How It Works · Data refreshed daily, snapshot 2026-10-07.