WeClawArena - Clinical: leaderboard
Metric: Task success rate (%; deterministic check of the final multi-workspace state over all five variants of each of the 10 Clinical base tasks, one no-attacker control and four attack-vector variants; each model drives every owner agent in the Dockerized OpenClaw runtime; rows with missing task-success fields count as failures). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Sonnet 4.5 (OpenClaw) | 40 |
| 2 | DeepSeek V3.2 (OpenClaw) | 40 |
| 3 | Kimi K2.5 (OpenClaw) | 40 |
| 4 | Claude Opus 4.1 (OpenClaw) | 38 |
| 5 | Claude Opus 4.7 (OpenClaw) | 36 |
| 6 | Qwen3 235B (OpenClaw) | 34 |
| 7 | Qwen3 32B (OpenClaw) | 28 |
| 8 | Kimi K2 Thinking (OpenClaw) | 26 |
Interactive version: theaggregate.ai/benchmark?slug=weclawarena-clinical · How It Works · Data refreshed daily, snapshot 2026-09-26.