MCP-Universe (LLM w/ ReAct) — leaderboard

MCP-Universe benchmarks LLMs on application-centric multi-step workflows: location navigation, repository management, financial analysis, 3D design, browser automation, and web search. ReAct prompting track.

Metric: Avg Evaluator Score. Source: mcp-universe.github.io. Status: years away from saturation. 28 models tracked.

Top models

#ModelScore
1GPT-5 (High)62.82
2GPT-5 (Medium)60.23
3Claude Sonnet 4 (Thinking)51.74
4Claude Sonnet 450.61
5Claude Opus 4.149.14
6Grok 449.01
7Grok 4 Fast48.95
8Grok Code Fast 144.72
9GLM-4.643.94
10DeepSeek V3.143.23
11DeepSeek V3.1 Terminus42.68
12DeepSeek V3.2 Exp41.92
13Qwen 3 Coder 480B A35B Instruct41.39
14GPT-4.141.32
15Kimi K2 090541.28

Interactive version: theaggregate.ai/benchmark?slug=mcp-universe-llm-w-react · How the rankings work · Data refreshed daily, snapshot 2026-07-22.