WeClawArena - Trading: leaderboard

Metric: Task success rate (%; deterministic check of the final multi-workspace state over all five variants of each of the 8 Trading base tasks, one no-attacker control and four attack-vector variants; each model drives every owner agent in the Dockerized OpenClaw runtime; rows with missing task-success fields count as failures). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 8 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.5 (OpenClaw)35
2Qwen3 32B (OpenClaw)35
3Claude Opus 4.7 (OpenClaw)30
4Kimi K2 Thinking (OpenClaw)30
5Qwen3 235B (OpenClaw)30
6Kimi K2.5 (OpenClaw)27.5
7Claude Opus 4.1 (OpenClaw)25
8DeepSeek V3.2 (OpenClaw)20

Interactive version: theaggregate.ai/benchmark?slug=weclawarena-trading · How It Works · Data refreshed daily, snapshot 2026-09-26.