MCP-Universe: leaderboard

Benchmark for LLMs and agents using real-world Model Context Protocol servers across location navigation, repository management, finance, 3D design, browser automation, and web search tasks.

Metric: Overall Success Rate (self-reported). Source: benchmarklist.com. Status: saturated. 28 models tracked.

Top models

#ModelScore
1GPT-5 (High)44.16
2GPT-5 (Medium)43.72
3Grok 433.33
4Claude Sonnet 4 (Thinking)30.3
5Claude Sonnet 429.44
6Claude Opus 4.129.44
7Grok 4 Fast27.27
8Grok Code Fast 126.41
9O3 (Medium)26.41
10GLM-4.625.97
11O4 Mini (Medium)25.97
12GLM-4.524.68
13Claude 3.7 Sonnet24.24
14Qwen 3 Coder 480B A35B Instruct22.94
15Gemini 2.5 Pro22.08

Interactive version: theaggregate.ai/benchmark?slug=mcp-universe · How It Works · Data refreshed daily, snapshot 2026-09-05.