WeClawArena - Trading: leaderboard
Metric: Task success rate (%; deterministic check of the final multi-workspace state over all five variants of each of the 8 Trading base tasks, one no-attacker control and four attack-vector variants; each model drives every owner agent in the Dockerized OpenClaw runtime; rows with missing task-success fields count as failures). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Sonnet 4.5 (OpenClaw) | 35 |
| 2 | Qwen3 32B (OpenClaw) | 35 |
| 3 | Claude Opus 4.7 (OpenClaw) | 30 |
| 4 | Kimi K2 Thinking (OpenClaw) | 30 |
| 5 | Qwen3 235B (OpenClaw) | 30 |
| 6 | Kimi K2.5 (OpenClaw) | 27.5 |
| 7 | Claude Opus 4.1 (OpenClaw) | 25 |
| 8 | DeepSeek V3.2 (OpenClaw) | 20 |
Interactive version: theaggregate.ai/benchmark?slug=weclawarena-trading · How It Works · Data refreshed daily, snapshot 2026-09-26.