MCP-Universe (Agent Track) — leaderboard
MCP-Universe agent track benchmarks autonomous agents (not raw LLMs) on complex enterprise workflows including 3D design and financial analysis.
Metric: Overall Success Rate. Source: mcp-universe.github.io. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | ReAct (GPT-5-Medium) | 43.72 |
| 2 | OpenAI Agent SDK (o3-Medium) | 31.6 |
| 3 | HarmonyReAct (GPT-OSS-120B) | 31.17 |
| 4 | ReAct (Claude-4.0-Sonnet) | 29.44 |
| 5 | Cursor Agent 1.4.5 (Claude-4.0-Sonnet-Thinking) | 26.41 |
| 6 | ReAct (o3-Medium) | 26.41 |
| 7 | HarmonyReAct (GPT-OSS-20B) | 24.24 |
| 8 | OpenAI Agent SDK (GPT-OSS-120B) | 11.26 |
Interactive version: theaggregate.ai/benchmark?slug=mcp-universe-agent-track · How the rankings work · Data refreshed daily, snapshot 2026-07-22.