MCP-Universe (LLM w/ ReAct) — leaderboard
MCP-Universe benchmarks LLMs on application-centric multi-step workflows: location navigation, repository management, financial analysis, 3D design, browser automation, and web search. ReAct prompting track.
Metric: Avg Evaluator Score. Source: mcp-universe.github.io. Status: years away from saturation. 28 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5 (High) | 62.82 |
| 2 | GPT-5 (Medium) | 60.23 |
| 3 | Claude Sonnet 4 (Thinking) | 51.74 |
| 4 | Claude Sonnet 4 | 50.61 |
| 5 | Claude Opus 4.1 | 49.14 |
| 6 | Grok 4 | 49.01 |
| 7 | Grok 4 Fast | 48.95 |
| 8 | Grok Code Fast 1 | 44.72 |
| 9 | GLM-4.6 | 43.94 |
| 10 | DeepSeek V3.1 | 43.23 |
| 11 | DeepSeek V3.1 Terminus | 42.68 |
| 12 | DeepSeek V3.2 Exp | 41.92 |
| 13 | Qwen 3 Coder 480B A35B Instruct | 41.39 |
| 14 | GPT-4.1 | 41.32 |
| 15 | Kimi K2 0905 | 41.28 |
Interactive version: theaggregate.ai/benchmark?slug=mcp-universe-llm-w-react · How the rankings work · Data refreshed daily, snapshot 2026-07-22.