MCP-Persona - Execution Accuracy: leaderboard

Metric: Execution accuracy (%): executor-verified correctness of the human-specified execution steps (final answers and tool-call parameters, CRUD state changes checked in the sandbox), averaged over tasks; 173 human-verified personal-application tasks on 12 simulated MCP servers (Lark, Slack, WeCom, Notion, Obsidian, Rednote, Reddit, Instagram, email and search tools), up to 20 tool-calling rounds, GPT-4o judge; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 13 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.541.5
2GPT-541.45
3Claude Opus 4.136.77
4O4 Mini34.73
5O330.27
6DeepSeek V3 (0324)27.52
7Grok 422.49
8Qwen 3 235B A22B 2507 Instruct21.83
9Qwen3 Coder20.14
10GPT-4o (2024-11-20)20.02
11Gemini 2.5 Pro13.38
12Gemini 3 Pro5.78

Interactive version: theaggregate.ai/benchmark?slug=mcp-persona-execution-accuracy · How It Works · Data refreshed daily, snapshot 2026-09-29.