ACE-Bench (Azure SDK, MCP Documentation): leaderboard

Metric: Strict pass rate (%) on ACE-Bench (353 Azure SDK coding tasks in Java, JavaScript/TypeScript, C# and Python built from official documentation examples; a task passes only if every atomic criterion holds: regex checks of required API usage and reference-grounded LLM-judge checks); same lightweight mcp-use coding agent for every model, augmented: the agent may query the Microsoft Learn MCP documentation server before answering; higher is better. Source: arxiv.org. Saturation forecast: Around September 2027. 11 models tracked.

Top models

#ModelScore
1Grok 468.7
2Claude Opus 4.565.3
3Claude Sonnet 4.563.7
4GPT-5.163.2
5Grok Code Fast 158.5
6Claude Haiku 4.558
7Claude Opus 4.153.6
8GPT-553.5
9GPT-4.151.1
10Grok 4 Fast (Non-reasoning)50.9
11GPT-5 Mini49.6

Interactive version: theaggregate.ai/benchmark?slug=ace-bench-azure-sdk-mcp-documentation · How It Works · Data refreshed daily, snapshot 2026-10-07.