MCP-Persona - Email: leaderboard

Metric: Checkpoint accuracy (%): mean LLM-judged score (0, 0.5 or 1) of the sub-task checkpoints in a task, averaged over tasks, single-server email tasks; 173 human-verified personal-application tasks on 12 simulated MCP servers (Lark, Slack, WeCom, Notion, Obsidian, Rednote, Reddit, Instagram, email and search tools), up to 20 tool-calling rounds, GPT-4o judge; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 13 models tracked.

Top models

#ModelScore
1O4 Mini53.83
2GPT-547.17
3Claude Sonnet 4.543.63
4O341.08
5Gemini 3 Pro36.92
6DeepSeek V3 (0324)30.79
7Gemini 2.5 Pro20.92
8Qwen 3 235B A22B 2507 Instruct13.71
9GPT-4o (2024-11-20)12.57
10Claude Opus 4.19.71
11Qwen3 Coder8.29
12Grok 45.71

Interactive version: theaggregate.ai/benchmark?slug=mcp-persona-email · How It Works · Data refreshed daily, snapshot 2026-09-29.