WeClawArena - SWE-Workspace: leaderboard

Metric: Task success rate (%; deterministic check of the final multi-workspace state over all five variants of each of the 50 SWE-Workspace base tasks, one no-attacker control and four attack-vector variants; each model drives every owner agent in the Dockerized OpenClaw runtime; rows with missing task-success fields count as failures). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 8 models tracked.

Top models

#ModelScore
1Claude Opus 4.7 (OpenClaw)34
2Claude Opus 4.1 (OpenClaw)24
3Kimi K2.5 (OpenClaw)15
4DeepSeek V3.2 (OpenClaw)14.8
5Claude Sonnet 4.5 (OpenClaw)8
6Kimi K2 Thinking (OpenClaw)6
7Qwen3 235B (OpenClaw)2
8Qwen3 32B (OpenClaw)1.6

Interactive version: theaggregate.ai/benchmark?slug=weclawarena-swe-workspace · How It Works · Data refreshed daily, snapshot 2026-09-26.