ClawArena-Team: leaderboard

Metric: Subagent-Management Score (%): task correctness multiplied by a least-privilege and modality-routing management factor in [0, 1]; 41 multi-turn multimodal scenarios with 258 evaluation rounds; a text-only main agent with partial workspace access creates, empowers and schedules a fixed locally served pool of LLM, VLM and omni subagents; execution-based checks, no LLM judge; higher is better. Source: arxiv.org. Saturation forecast: Around August 2027. 12 models tracked.

Top models

#ModelScore
1Claude Fable 560
2Gemini 3.5 Flash53.8
3GPT-5.551
4GLM-5.250.4
5GPT-5.449.7
6Gemini 3.1 Pro (Preview)49.6
7Kimi K2.649
8Claude Sonnet 4.648.5
9DeepSeek V4 Pro46.4
10Qwen 3.6 27B46.4
11Gemma 4 31B43.9
12GLM-4.7 Flash15.3

Interactive version: theaggregate.ai/benchmark?slug=clawarena-team · How It Works · Data refreshed daily, snapshot 2026-09-29.