OccuBench: leaderboard

Professional-task benchmark using simulated domain tool environments to evaluate LLM agents across occupation-specific workflows.

Metric: Completion rate (self-reported). Source: benchmarklist.com. Status: saturation imminent. 15 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)45.3
2GPT-5.242.6
3Qwen 3.5 Plus41
4Claude Opus 4.640.1
5MiniMax-M2.739.4
6GLM-5.135.7
7Kimi K2.533.9
8GPT-4o28.8
9Qwen 3 Max24.8
10Llama 4 Maverick22
11Gemma 3 27B7
12DeepSeek V3.21.9
13Gemma 3 12B1.5
14Qwen 3.5 4B0.2

Interactive version: theaggregate.ai/benchmark?slug=occubench · How It Works · Data refreshed daily, snapshot 2026-09-05.