ParamBench (3-shot): leaderboard

Metric: Exact match (%; share of the 293 held-out ParamBench test calls, 24 cloud-network APIs unseen in training, whose generated parameter object equals the gold object exactly; per-call protocol with the target API schema and the outputs of earlier calls, 3-shot prompting). Source: arxiv.org. Saturation forecast: Around July 2028. 9 models tracked.

Top models

#ModelScore
1GPT-5.441.3
2Qwen 3.6 Plus39.2
3DeepSeek V4 Pro38.9
4Claude Opus 4.734.5

Interactive version: theaggregate.ai/benchmark?slug=parambench-3-shot · How It Works · Data refreshed daily, snapshot 2026-09-29.