WebTestBench - Constraint: leaderboard

Metric: Defect-detection F1 (%) on latent logical-constraint test items, on 100 generated web applications, WebTester agents on Claude Code with Playwright MCP, predicted test items matched to the human gold checklist by a Qwen3.5-27B judge; higher is better. Source: arxiv.org. Saturation forecast: Not forecast. 10 models tracked.

Top models

#ModelScoreOverall rank
1MiMo-V2-Flash29.2#359
2Step 3.5 Flash27.9#262
3GPT-5.126.9#131
4GLM-526.9#137
5Qwen3 Coder Next23.8#321
6Claude Sonnet 4.522.5#138
7GPT-5.221.5#105
8Claude Opus 4.521.2#79
9GLM-4.720.5#185
10MiniMax-M2.115.8#283

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=webtestbench-constraint · How It Works · Data refreshed daily, snapshot 2026-10-11.