APEX-Agents-AA — leaderboard

Artificial Analysis implementation of APEX-Agents using the Stirrup agent harness for long-horizon, cross-application professional-services tasks.

Metric: Pass@1 (self-reported). Source: benchmarklist.com. Status: saturation imminent. 24 models tracked.

Top models

#ModelScore
1Gemini 3.5 Flash47.1
2GPT-5.537.7
3GPT-5.433.3
4Claude Opus 4.6 (Max)33
5Gemini 3.1 Pro (Preview)32
6GPT-5.4 Mini28.2
7Claude Sonnet 4.6 (Max)28
8Gemini 3 Flash (Preview)27.7
9GPT-5.4 Nano24.9
10DeepSeek V4 Pro (Max)24.3
11Qwen 3.7 Plus22.4
12Grok 4.317
13Qwen 3.5 397B A17B15.3
14DeepSeek V3.214.5
15GLM-514.5

Interactive version: theaggregate.ai/benchmark?slug=apex-agents-aa · How the rankings work · Data refreshed daily, snapshot 2026-07-22.