BlueBench - Entity Extraction — leaderboard

Metric: Score (%). Source: huggingface.co. 18 models tracked.

Top models

#ModelScore
1O179.22
2O4 Mini75.47
3Mistral Medium 374.21
4GPT-4o73.62
5GPT-4.1 Mini73.17
6O3 Mini72.86
7GPT-4.172.39
8Llama 3.3 70B Instruct70.18
9Mistral Large61.18
10GPT-4.1 Nano61.11
11Llama 3.2 3B Instruct45.26
12Pixtral-12B22
13Llama 3.2 1B Instruct20.69

Interactive version: theaggregate.ai/benchmark?slug=bluebench-entity-extraction · How the rankings work · Data refreshed daily, snapshot 2026-07-22.