ClawArena-Team - Tool-Permission Precision: leaderboard

Metric: Tool-permission precision (%): per subagent, the share of granted tool types actually used, averaged; 41 multi-turn multimodal scenarios with 258 evaluation rounds; a text-only main agent with partial workspace access creates, empowers and schedules a fixed locally served pool of LLM, VLM and omni subagents; execution-based checks, no LLM judge; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 12 models tracked.

Top models

#ModelScore
1GPT-5.579.9
2Gemma 4 31B79
3Claude Sonnet 4.677.3
4GPT-5.477.3
5Kimi K2.676.8
6Gemini 3.1 Pro (Preview)76.4
7Claude Fable 576.4
8DeepSeek V4 Pro73.9
9Qwen 3.6 27B72.5
10Gemini 3.5 Flash69.8
11GLM-5.267.4
12GLM-4.7 Flash22.4

Interactive version: theaggregate.ai/benchmark?slug=clawarena-team-tool-permission-precision · How It Works · Data refreshed daily, snapshot 2026-09-29.