ClawArena (Claude Code): leaderboard

Metric: Composite Reliability Score (CRS, 0-100) on ClawArena (12 multi-turn scenarios, 337 evaluation rounds with staged updates; multi-choice and shell executable-check questions): the mean of the task completion rate and Robustness (success cohesion times failure dispersion), macro-averaged over scenarios; agent run in the Claude Code framework; higher is better. Source: arxiv.org. Saturation forecast: Around 2031. 6 models tracked.

Top models

#ModelScore
1Claude Opus 4.766.31
2Claude Sonnet 4.662.16
3Claude Haiku 4.560.93
4Kimi K2.559.75
5GPT-5.147.57
6GPT-5.544.5

Interactive version: theaggregate.ai/benchmark?slug=clawarena-claude-code · How It Works · Data refreshed daily, snapshot 2026-10-07.