WeClawArena - Clinical: leaderboard

Metric: Task success rate (%; deterministic check of the final multi-workspace state over all five variants of each of the 10 Clinical base tasks, one no-attacker control and four attack-vector variants; each model drives every owner agent in the Dockerized OpenClaw runtime; rows with missing task-success fields count as failures). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 8 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.5 (OpenClaw)40
2DeepSeek V3.2 (OpenClaw)40
3Kimi K2.5 (OpenClaw)40
4Claude Opus 4.1 (OpenClaw)38
5Claude Opus 4.7 (OpenClaw)36
6Qwen3 235B (OpenClaw)34
7Qwen3 32B (OpenClaw)28
8Kimi K2 Thinking (OpenClaw)26

Interactive version: theaggregate.ai/benchmark?slug=weclawarena-clinical · How It Works · Data refreshed daily, snapshot 2026-09-26.