WeClawArena - Travel: leaderboard

Metric: Task success rate (%; deterministic check of the final multi-workspace state over all five variants of each of the 20 Travel base tasks, one no-attacker control and four attack-vector variants; each model drives every owner agent in the Dockerized OpenClaw runtime; rows with missing task-success fields count as failures). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 8 models tracked.

Top models

#ModelScore
1Claude Opus 4.7 (OpenClaw)83
2Claude Sonnet 4.5 (OpenClaw)58
3Kimi K2.5 (OpenClaw)58
4Qwen3 235B (OpenClaw)47
5Claude Opus 4.1 (OpenClaw)46
6DeepSeek V3.2 (OpenClaw)26
7Qwen3 32B (OpenClaw)20
8Kimi K2 Thinking (OpenClaw)11

Interactive version: theaggregate.ai/benchmark?slug=weclawarena-travel · How It Works · Data refreshed daily, snapshot 2026-09-26.