BlueBench - RAG General — leaderboard

Metric: Score (%). Source: huggingface.co. 18 models tracked.

Top models

#ModelScore
1GPT-4.1 Mini54.58
2Mistral Medium 352.57
3Mistral Large51.08
4GPT-4.1 Nano50.88
5Llama 3.3 70B Instruct49.02
6GPT-4o48.9
7GPT-4.148.46
8Pixtral-12B47.85
9Llama 3.2 3B Instruct47.81
10Llama 3.2 1B Instruct44.7
11O3 Mini44.06
12O141.24
13O4 Mini40.32

Interactive version: theaggregate.ai/benchmark?slug=bluebench-rag-general · How the rankings work · Data refreshed daily, snapshot 2026-07-22.