WebTestBench: leaderboard

Metric: Defect-detection F1 (%), real defects (gold Fail items) as the positive class, end to end from the agent's own checklist, on 100 generated web applications, WebTester agents on Claude Code with Playwright MCP, predicted test items matched to the human gold checklist by a Qwen3.5-27B judge; higher is better. Source: arxiv.org. Saturation forecast: Around 2031. 10 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5.126.4#131
2MiMo-V2-Flash25.1#359
3Step 3.5 Flash23.4#262
4GPT-5.222.9#105
5Claude Sonnet 4.521.9#138
6Claude Opus 4.520.2#79
7GLM-519#137
8GLM-4.718.1#185
9Qwen3 Coder Next17.3#321
10MiniMax-M2.115.2#283

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=webtestbench · How It Works · Data refreshed daily, snapshot 2026-10-11.