Skip to content
The Aggregate

Unified LLM rankings, daily

Models What's New Trends Benchmarks Skill Maps Task Explorer
How It Works Guesswork Metrics
The Aggregate
Loading data...

CodeRouterBench — leaderboard

Metric: Mean Task Score (%). Source: huggingface.co. 8 models tracked.

Top models

#ModelScore
1Claude Opus 4.643.83
2GPT-5.442.21
3Claude Sonnet 4.641.31
4Qwen 3 Max39.67
5GLM-538.03
6Qwen 3.5 Plus37.16
7Kimi K2.536.66
8MiniMax-M2.735.62

Interactive version: theaggregate.ai/benchmark?slug=coderouterbench · How It Works · Data refreshed daily, snapshot 2026-07-27.

Built by Mikhail Doroshenko — AI researcher, co-author of Humanity’s Last Exam. Independent project; no affiliation with any AI lab.