WeClawArena - Bargaining: leaderboard
Metric: Task success rate (%; deterministic check of the final multi-workspace state over all five variants of each of the 24 Bargaining base tasks, one no-attacker control and four attack-vector variants; each model drives every owner agent in the Dockerized OpenClaw runtime; rows with missing task-success fields count as failures). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Sonnet 4.5 (OpenClaw) | 68.3 |
| 2 | Claude Opus 4.7 (OpenClaw) | 63.3 |
| 3 | Kimi K2.5 (OpenClaw) | 47.5 |
| 4 | Qwen3 235B (OpenClaw) | 25.8 |
| 5 | Claude Opus 4.1 (OpenClaw) | 22.5 |
| 6 | DeepSeek V3.2 (OpenClaw) | 20 |
| 7 | Kimi K2 Thinking (OpenClaw) | 19.2 |
| 8 | Qwen3 32B (OpenClaw) | 11.7 |
Interactive version: theaggregate.ai/benchmark?slug=weclawarena-bargaining · How It Works · Data refreshed daily, snapshot 2026-09-26.