ClawArena-Team - Workspace-Permission Precision: leaderboard

Metric: Workspace-permission precision (%): per subagent, files accessed divided by files granted, averaged; 41 multi-turn multimodal scenarios with 258 evaluation rounds; a text-only main agent with partial workspace access creates, empowers and schedules a fixed locally served pool of LLM, VLM and omni subagents; execution-based checks, no LLM judge; higher is better. Source: arxiv.org. Saturation forecast: Around April 2028. 12 models tracked.

Top models

#ModelScore
1Claude Fable 549.2
2GPT-5.547.5
3DeepSeek V4 Pro46.3
4Gemini 3.5 Flash45.7
5GLM-5.244.4
6Claude Sonnet 4.643.2
7Kimi K2.641.9
8GPT-5.440.8
9Qwen 3.6 27B40.7
10Gemma 4 31B37.9
11Gemini 3.1 Pro (Preview)36.6
12GLM-4.7 Flash11.7

Interactive version: theaggregate.ai/benchmark?slug=clawarena-team-workspace-permission-precision · How It Works · Data refreshed daily, snapshot 2026-09-29.