IndustryBench-MIPU - Single-Image: leaderboard

Metric: Image-level F1 (%) of single-image extraction (each image processed independently), structured attribute value extraction from industrial product images (specification tables, nameplates, technical drawings): property names matched exactly and values by rule-based normalization then a Qwen 3.6 Plus semantic judge; full extraction prompt, thinking enabled where supported; higher is better. Source: arxiv.org. Saturation forecast: Around April 2028. 9 models tracked.

Top models

#ModelScore
1Qwen 3.5 Plus (Thinking)81.3
2Gemini 3.1 Pro (Preview)77.1
3Qwen 3.5 397B A17B76
4Qwen 3.5 27B71.5
5Kimi K2.5 (Thinking)70.9
6Qwen 3.5 122B A10B70.5
7Claude Opus 4.6 (Thinking)69.4
8Qwen 3.5 35B A3B68.7
9GPT-5.4 (Thinking)66.2

Interactive version: theaggregate.ai/benchmark?slug=industrybench-mipu-single-image · How It Works · Data refreshed daily, snapshot 2026-09-29.