DynamicMCPBench: leaderboard

Metric: Pass^3 (%; share of the 750 balanced tasks from 15 tool-use categories over 121 MCP servers solved in all three independent attempts; effect checkpoints scored deterministically under replay, no LLM judge; tools needed plus eight distractors offered). Source: arxiv.org. Saturation forecast: Around May 2028. 24 models tracked.

Top models

#ModelScore
1Qwen 3.7 Max51.2
2GLM-5.150.3
3Qwen 3.6 35B A3B48.5
4DeepSeek V4 Pro46.4
5Gemma 4 31B42.5
6MiniMax-M342.4
7Claude Haiku 4.541.1
8Kimi K2.640.7
9Grok 4.331.6
10Qwen 3.5 4B27.3
11GPT-5.4 Mini25.7
12Qwen 3 8B22.1
13Gemma 4 E4B22.1
14Qwen 2.5 7B21.1
15Gemma 4 E2B20.1

Interactive version: theaggregate.ai/benchmark?slug=dynamicmcpbench · How It Works · Data refreshed daily, snapshot 2026-09-29.