IndustryBench (Thinking): leaderboard

Metric: Final (SV) score (out of 3): mean 0-3 rubric score from a Qwen3-Max judge against the reference answer, set to 0 for any answer with a safety violation against the source standard; closed-book, zero-shot, thinking mode enabled, only the final answer judged, over the 2,049 Chinese industrial procurement QA items of IndustryBench grounded in Chinese national standards (GB/T); higher is better. Source: arxiv.org. Saturation forecast: Around May 2027. 13 models tracked.

Top models

#ModelScore
1Claude Opus 4.6 (Thinking)2.03
2GPT-5.4 (Thinking)1.98
3Gemini 3.1 Pro (Preview)1.97
4Qwen 3.6 Plus (Thinking)1.89
5Qwen 3.5 397B A17B1.81
6Qwen 3.5 Plus (Thinking)1.79
7Qwen 3 Max (Thinking)1.75
8GLM-51.72
9Qwen 3.5 122B A10B1.71
10Kimi K2.51.68
11Qwen 3.5 27B1.65
12Qwen 3.5 35B A3B1.64
13MiniMax-M2.51.42

Interactive version: theaggregate.ai/benchmark?slug=industrybench-thinking · How It Works · Data refreshed daily, snapshot 2026-10-07.