PBT-Bench (Baseline Prompt): leaderboard

Metric: Bug recall (%) on the 100 problems under the open-ended baseline prompt (write tests that expose the injected bugs, no property-based testing scaffolding), OpenHands agent scaffold, mean of three runs, each problem weighted equally; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 8 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.676.7
2GLM-5.166.6
3DeepSeek V3.264.2
4Gemini 3 Flash56.2
5Qwen 3.6 Plus53.5
6Grok 4.1 Fast50.1
7Step 3.5 Flash37.8

Interactive version: theaggregate.ai/benchmark?slug=pbt-bench-baseline-prompt · How It Works · Data refreshed daily, snapshot 2026-10-07.