WebTestBench - Interaction: leaderboard

Metric: Defect-detection F1 (%) on interaction test items, on 100 generated web applications, WebTester agents on Claude Code with Playwright MCP, predicted test items matched to the human gold checklist by a Qwen3.5-27B judge; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 10 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5.223.2#105
2GPT-5.122#131
3Step 3.5 Flash21.2#262
4GLM-520.9#137
5MiMo-V2-Flash20.3#359
6Claude Sonnet 4.519.9#138
7MiniMax-M2.119.9#283
8GLM-4.717.2#185
9Claude Opus 4.514.7#79
10Qwen3 Coder Next11.4#321

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=webtestbench-interaction · How It Works · Data refreshed daily, snapshot 2026-10-11.