BLXBench: leaderboard

Community benchmark runner and public leaderboard for AI model performance across coding, debugging, reasoning, hallucination, refactoring, security, and speed slices.

Metric: Score (self-reported). Source: benchmarklist.com. 25 models tracked.

Top models

#ModelScore
1Grok 4.385.5
2Claude Opus 4.784.8
3Qwen 3.6 Flash82.8
4Grok 4.2079.1
5Ling-2.6-1T75.3
6Mistral Small 475.2
7Claude Opus 4.671.1
8GPT-5.565.9
9Granite 4.1 8B64.6
10DeepSeek V4 Flash48.3
11GPT-5.5 Pro44.5
12MiMo-V2.5-Pro44
13Qwen 3.6 35B A3B33.9
14Qwen 3.6 27B29
15MiMo-V2.515.7

Interactive version: theaggregate.ai/benchmark?slug=blxbench · How It Works · Data refreshed daily, snapshot 2026-09-05.