APEX-Agents — leaderboard

The AI Productivity Index for Agents (APEX-Agents) measures whether frontier AI agents can execute long-horizon, cross-application tasks across three jobs in professional services.

Metric: Mean Score (ReAct) (self-reported). Source: benchmarklist.com. Status: saturation imminent. 40 models tracked.

Top models

#ModelScore
1Gemini 3.5 Flash66.1
2Claude Opus 4.859.4
3GPT-5.5 (xHigh)53.9
4GPT-5.4 (xHigh)52.7
5Claude Opus 4.750.6
6GPT-5.2 (xHigh)48.4
7Claude Opus 4.6 (Max)48.4
8Gemini 3.1 Pro (Preview) (High)48.2
9GPT-5.3 Codex (High)46.9
10GPT-5.2 Codex (High)42.2
11Claude Sonnet 4.6 (High)40.7
12Gemini 3 Flash (Preview) (High)39.5
13GPT-5.4 Mini (xHigh)37.5
14GPT-5.1 Codex (High)34.9
15GPT-5 Codex (High)34.8

Interactive version: theaggregate.ai/benchmark?slug=apex-agents · How the rankings work · Data refreshed daily, snapshot 2026-07-22.