APEX-Agents: leaderboard

The AI Productivity Index for Agents (APEX-Agents) measures whether frontier AI agents can execute long-horizon, cross-application tasks across three jobs in professional services.

Metric: Mean Score (ReAct) (self-reported). Source: benchmarklist.com. Status: saturation imminent. 45 models tracked.

Top models

#ModelScore
1Gemini 3.5 Flash66.1
2GPT-5.4 (xHigh)52.7
3Claude Opus 4.7 (Max)50.6
4GPT-5.2 (xHigh)48.4
5Claude Opus 4.6 (Max)48.4
6Gemini 3.1 Pro (Preview) (High)48.2
7GPT-5.3 Codex (High)46.9
8Claude Opus 4.6 (High)45.6
9Kimi K3 (Max)41
10Claude Sonnet 4.6 (High)40.7
11GPT-5.6 Sol (Max)39.9
12Gemini 3 Flash (High)39.5
13Claude Opus 4.8 (Max)39.4
14GPT-5.5 (xHigh)38.5
15GPT-5.4 Mini (xHigh)37.5

Interactive version: theaggregate.ai/benchmark?slug=apex-agents · How It Works · Data refreshed daily, snapshot 2026-09-05.