WebTestBench: leaderboard
Metric: Defect-detection F1 (%), real defects (gold Fail items) as the positive class, end to end from the agent's own checklist, on 100 generated web applications, WebTester agents on Claude Code with Playwright MCP, predicted test items matched to the human gold checklist by a Qwen3.5-27B judge; higher is better. Source: arxiv.org. Saturation forecast: Around 2031. 10 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | GPT-5.1 | 26.4 | #131 |
| 2 | MiMo-V2-Flash | 25.1 | #359 |
| 3 | Step 3.5 Flash | 23.4 | #262 |
| 4 | GPT-5.2 | 22.9 | #105 |
| 5 | Claude Sonnet 4.5 | 21.9 | #138 |
| 6 | Claude Opus 4.5 | 20.2 | #79 |
| 7 | GLM-5 | 19 | #137 |
| 8 | GLM-4.7 | 18.1 | #185 |
| 9 | Qwen3 Coder Next | 17.3 | #321 |
| 10 | MiniMax-M2.1 | 15.2 | #283 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=webtestbench · How It Works · Data refreshed daily, snapshot 2026-10-11.