ClawArena-Team - Task Completion: leaderboard

Metric: Task completion rate (%): mean pass rate over user questions; 41 multi-turn multimodal scenarios with 258 evaluation rounds; a text-only main agent with partial workspace access creates, empowers and schedules a fixed locally served pool of LLM, VLM and omni subagents; execution-based checks, no LLM judge; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 12 models tracked.

Top models

#ModelScore
1Claude Fable 574.4
2Gemini 3.5 Flash69.8
3GLM-5.266.3
4Gemini 3.1 Pro (Preview)65.5
5Kimi K2.664.3
6GPT-5.563.6
7GPT-5.463.2
8Claude Sonnet 4.661.6
9Qwen 3.6 27B60.9
10DeepSeek V4 Pro58.9
11Gemma 4 31B56.6
12GLM-4.7 Flash34.5

Interactive version: theaggregate.ai/benchmark?slug=clawarena-team-task-completion · How It Works · Data refreshed daily, snapshot 2026-09-29.