OccuBench - Industrial & Engineering: leaderboard

Metric: Completion rate (%) on the Industrial & Engineering industry's tasks in the clean environment (E0); 382 tool-use task instances from 100 professional scenarios, the environment simulated by Gemini-3-Flash-Preview and each trajectory passed or failed by a rubric verifier; thinking mode on, high reasoning effort where configurable; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 15 models tracked.

Top models

#ModelScore
1GPT-5.2 (High)85
2Gemini 3.1 Pro (Preview)73
3Claude Opus 4.6 (High)73
4Claude Sonnet 4.5 (Thinking)71
5Qwen 3.5 Plus (Thinking)71
6DeepSeek V3.2 (Thinking)69
7Claude Opus 4.5 (Thinking)65
8Claude Sonnet 4.6 (Thinking)64
9Kimi K2.5 (Thinking)62
10MiniMax-M2.760
11Claude Opus 4 (Thinking)58
12GLM-5 (Thinking)56
13Qwen 3.5 Flash (Thinking)53
14Claude Sonnet 4 (Thinking)51

Interactive version: theaggregate.ai/benchmark?slug=occubench-industrial-engineering · How It Works · Data refreshed daily, snapshot 2026-10-07.