BlueBench - Bias — leaderboard

Metric: Score (%). Source: huggingface.co. 18 models tracked.

Top models

#ModelScore
1GPT-4.197.98
2O3 Mini97.98
3Mistral Large96.97
4Llama 3.3 70B Instruct95.96
5GPT-4.1 Mini94.95
6O194.95
7O4 Mini94.95
8GPT-4o93.94
9Mistral Medium 393.94
10GPT-4.1 Nano81.82
11Pixtral-12B72.73
12Llama 3.2 3B Instruct60.61
13Llama 3.2 1B Instruct48.48

Interactive version: theaggregate.ai/benchmark?slug=bluebench-bias · How the rankings work · Data refreshed daily, snapshot 2026-07-22.