ORAgentBench — leaderboard
Metric: Pass Rate All (self-reported). Source: benchmarklist.com. 14 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.4 | 35.51 |
| 2 | Claude Opus 4.6 | 34.58 |
| 3 | GPT-5.3 Codex | 32.71 |
| 4 | DeepSeek V4 Pro | 27.1 |
| 5 | GLM-5.1 | 26.17 |
| 6 | GPT-5.4 Mini | 26.17 |
| 7 | DeepSeek V4 Flash | 26.17 |
| 8 | Claude Sonnet 4.6 | 25.23 |
| 9 | Qwen 3.6 Plus | 22.43 |
| 10 | GLM-5 | 22.43 |
| 11 | MiMo-V2.5-Pro | 22.43 |
| 12 | Qwen 3.5 Plus (2026-04-20) | 20.56 |
| 13 | MiniMax-M2.7 | 14.02 |
Interactive version: theaggregate.ai/benchmark?slug=oragentbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.