WeClawArena - Bidding: leaderboard

Metric: Task success rate (%; deterministic check of the final multi-workspace state over all five variants of each of the 12 Bidding base tasks, one no-attacker control and four attack-vector variants; each model drives every owner agent in the Dockerized OpenClaw runtime; rows with missing task-success fields count as failures). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 8 models tracked.

Top models

#ModelScore
1Claude Opus 4.7 (OpenClaw)55
2Claude Sonnet 4.5 (OpenClaw)51.7
3Kimi K2.5 (OpenClaw)40
4Claude Opus 4.1 (OpenClaw)6.7
5DeepSeek V3.2 (OpenClaw)5
6Kimi K2 Thinking (OpenClaw)3.3
7Qwen3 235B (OpenClaw)3.3
8Qwen3 32B (OpenClaw)1.7

Interactive version: theaggregate.ai/benchmark?slug=weclawarena-bidding · How It Works · Data refreshed daily, snapshot 2026-09-26.