ParamBench (0-shot): leaderboard

Metric: Exact match (%; share of the 293 held-out ParamBench test calls, 24 cloud-network APIs unseen in training, whose generated parameter object equals the gold object exactly; per-call protocol with the target API schema and the outputs of earlier calls, 0-shot prompting). Source: arxiv.org. Saturation forecast: Around 2029. 9 models tracked.

Top models

#ModelScore
1Qwen 3.6 Plus39.2
2DeepSeek V4 Pro36.5
3GPT-5.432.4
4Claude Opus 4.729.4

Interactive version: theaggregate.ai/benchmark?slug=parambench-0-shot · How It Works · Data refreshed daily, snapshot 2026-09-29.