MCP-Universe (LLM w/ Function Calls) — leaderboard
MCP-Universe benchmarks LLMs on application-centric multi-step workflows across 6 task categories. Function calling track from Salesforce AI Research.
Metric: Avg Evaluator Score. Source: mcp-universe.github.io. Status: saturation imminent. 21 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Pro (Preview) | 61.67 |
| 2 | Grok 4.1 Fast | 60.58 |
| 3 | GPT-5 (Medium) | 59.29 |
| 4 | Claude Sonnet 4.5 | 54.38 |
| 5 | Claude Sonnet 4 | 52.61 |
| 6 | Grok 4 Fast | 52.2 |
| 7 | Claude Sonnet 4 (Thinking) | 50.83 |
| 8 | Kimi K2 (Thinking) | 45.23 |
| 9 | Claude Haiku 4.5 | 44.63 |
| 10 | Grok Code Fast 1 | 43.89 |
| 11 | GPT-OSS-120B | 39.78 |
| 12 | GPT-4.1 | 39.58 |
| 13 | DeepSeek V3.1 | 38.88 |
| 14 | Kimi K2 0905 | 38.52 |
| 15 | GLM-4.5 | 37.58 |
Interactive version: theaggregate.ai/benchmark?slug=mcp-universe-llm-w-function-calls · How the rankings work · Data refreshed daily, snapshot 2026-07-22.