WebTestBench - Checklist Coverage: leaderboard
Metric: Checklist coverage (%): share of gold test items the agent's generated checklist covers, on 100 generated web applications, WebTester agents on Claude Code with Playwright MCP, predicted test items matched to the human gold checklist by a Qwen3.5-27B judge; higher is better. Source: arxiv.org. Saturation forecast: Around June 2028. 10 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Step 3.5 Flash | 66 | #262 |
| 2 | Claude Sonnet 4.5 | 63.7 | #138 |
| 3 | MiMo-V2-Flash | 63.5 | #359 |
| 4 | Claude Opus 4.5 | 63.2 | #79 |
| 5 | GPT-5.1 | 63.1 | #131 |
| 6 | GLM-5 | 63.1 | #137 |
| 7 | GLM-4.7 | 61.1 | #185 |
| 8 | GPT-5.2 | 61 | #105 |
| 9 | Qwen3 Coder Next | 60.4 | #321 |
| 10 | MiniMax-M2.1 | 60.1 | #283 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=webtestbench-checklist-coverage · How It Works · Data refreshed daily, snapshot 2026-10-11.