MCP-Universe — leaderboard

Benchmark for LLMs and agents using real-world Model Context Protocol servers across location navigation, repository management, finance, 3D design, browser automation, and web search tasks.

Metric: Overall Success Rate (self-reported). Source: benchmarklist.com. Status: saturation imminent. 27 models tracked.

Top models

#ModelScore
1GPT-5 (High)44.16
2Grok 433.33
3Claude Sonnet 4 (Thinking)30.3
4Claude Sonnet 429.44
5Claude Opus 4.129.44
6Grok 4 Fast27.27
7Grok Code Fast 126.41
8O3 (Medium)26.41
9GLM-4.625.97
10O4 Mini (Medium)25.97
11GLM-4.524.68
12Claude 3.7 Sonnet24.24
13Gemini 2.5 Pro22.08
14DeepSeek V3.122.08
15Gemini 2.5 Flash21.65

Interactive version: theaggregate.ai/benchmark?slug=mcp-universe · How the rankings work · Data refreshed daily, snapshot 2026-07-22.