IndustryBench-MIPU: leaderboard

Metric: Product-level F1 (%) of multi-image extraction (the model receives all valid images of a product and outputs product-level property-value pairs), structured attribute value extraction from industrial product images (specification tables, nameplates, technical drawings): property names matched exactly and values by rule-based normalization then a Qwen 3.6 Plus semantic judge; full extraction prompt, thinking enabled where supported; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 9 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)65.1
2Qwen 3.5 397B A17B62.7
3GPT-5.4 (Thinking)60.5
4Qwen 3.5 Plus (Thinking)59.9
5Claude Opus 4.6 (Thinking)57.2
6Kimi K2.5 (Thinking)56.7
7Qwen 3.5 27B55.8
8Qwen 3.5 122B A10B50.1
9Qwen 3.5 35B A3B20.6

Interactive version: theaggregate.ai/benchmark?slug=industrybench-mipu · How It Works · Data refreshed daily, snapshot 2026-09-29.