BLXBench — leaderboard

Community benchmark runner and public leaderboard for AI model performance across coding, debugging, reasoning, hallucination, refactoring, security, and speed slices.

Metric: Score (self-reported). Source: benchmarklist.com. 25 models tracked.

Top models

#ModelScore
1Grok 4.385.5
2Claude Opus 4.784.8
3Grok 4.2079.1
4Ling-2.6-1T75.3
5Claude Opus 4.671.1
6GPT-5.565.9
7Granite 4.1 8B64.6
8DeepSeek V4 Flash48.3
9GPT-5.5 Pro44.5
10MiMo-V2.5-Pro44
11Qwen 3.6 35B A3B33.9
12Qwen 3.6 27B29
13MiMo-V2.515.7
14DeepSeek V4 Pro15.2
15GLM-5.113.9

Interactive version: theaggregate.ai/benchmark?slug=blxbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.