WebTestBench - Content: leaderboard

Metric: Defect-detection F1 (%) on content test items, on 100 generated web applications, WebTester agents on Claude Code with Playwright MCP, predicted test items matched to the human gold checklist by a Qwen3.5-27B judge; higher is better. Source: arxiv.org. Saturation forecast: Not forecast. 10 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5.115.3#131
2MiniMax-M2.17.7#283
3MiMo-V2-Flash7.3#359
4Claude Opus 4.56.8#79
5GPT-5.26.2#105
6GLM-4.74.3#185
7Qwen3 Coder Next4.3#321
8GLM-53.4#137
9Step 3.5 Flash2.6#262
10Claude Sonnet 4.51.7#138

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=webtestbench-content · How It Works · Data refreshed daily, snapshot 2026-10-11.