SDR-Bench: leaderboard

Metric: Weighted coverage score (%) on 180 public customer success stories: a GPT-4o judge grades each ground-truth pitch point 0 to 5 for how well the agent's predicted pitch covers it, normalized to a percentage; agents search a frozen view of the web at the story's historical date; higher is better. Source: arxiv.org. Saturation forecast: Around July 2028. 6 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.655.8
2GPT-5.4 Mini44.63
3GPT-5.444.32
4GPT-4o Mini37.46
5Qwen 2.5 72B36.84
6GPT-4o35.42

Interactive version: theaggregate.ai/benchmark?slug=sdr-bench · How It Works · Data refreshed daily, snapshot 2026-09-29.