AJ-Bench — leaderboard

Metric: Overall Avg@3 (%). Source: aj-bench.github.io. 17 models tracked.

Top models

#ModelScore
1DeepSeek V3.277.34
2Gemini 3 Pro (Preview)75.05
3GPT-5 Mini (Low)72.41
4Claude Sonnet 4.569.77
5Grok 469.24
6Gemini 2.5 Pro68.31
7Claude Opus 4.566.79
8GLM-4.664.81
9Kimi K2 090564.33
10GPT-561.02
11Qwen 3 235B A22B58.47
12LongCat-Flash-Chat57.5
13GPT-5.153.5

Interactive version: theaggregate.ai/benchmark?slug=aj-bench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.