EnterpriseOps-Gym — leaderboard

Stateful enterprise operations benchmark for LLM agents performing long-horizon planning, tool use, and policy-governed workflows.

Metric: Task Success Rate (self-reported). Source: benchmarklist.com. Status: saturation imminent. 22 models tracked.

Top models

#ModelScore
1Claude Opus 4.644.6
2Claude Sonnet 4.640.4
3Claude Opus 4.537
4Gemini 3.1 Pro (Preview)36.6
5Gemini 3 Flash (Preview)31.7
6GPT-5.2 (High)31.3
7Claude Sonnet 4.530.5
8GPT-529.2
9Nemotron 3 Super27.3
10Kimi K2.526.2
11DeepSeek V3.2 (High)23.8
12MiniMax-M2.723
13GPT-OSS-120B (High)23
14GLM-522.2
15GPT-5 Mini20.6

Interactive version: theaggregate.ai/benchmark?slug=enterpriseops-gym · How the rankings work · Data refreshed daily, snapshot 2026-07-22.