BlueBench — leaderboard

Metric: Average Score (%). Source: huggingface.co. 18 models tracked.

Top models

#ModelScore
1GPT-4.164.61
2GPT-4o64.13
3GPT-4.1 Mini63.01
4Mistral Medium 360.5
5Llama 3.3 70B Instruct58.96
6O158.6
7GPT-4.1 Nano56.98
8Mistral Large55.49
9O4 Mini53.98
10O3 Mini53.85
11Pixtral-12B45.38
12Llama 3.2 3B Instruct44.27
13Llama 3.2 1B Instruct31.83

Interactive version: theaggregate.ai/benchmark?slug=bluebench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.