ProLLM - Entity Extraction — leaderboard

ProLLM benchmark evaluating LLMs on structured entity extraction from unstructured text. Measured by F1 score.

Metric: Score (%). Source: www.prollm.ai. Status: saturated. 48 models tracked.

Top models

#ModelScore
1Grok 385.5
2GPT-4o85.4
3Mistral Small 385.3
4MiniMax-Text-0185.1
5O184.8
6Nova Pro84.5
7GPT-5 Nano84.2
8Gemma 2 9B (IT)83.9
9DeepSeek V383.9
10Gemini 2.0 Flash83.8
11GPT-4o Mini83.6
12Kimi K283.3
13O3 Mini (Medium)83.1
14Mistral Small 3.182.8
15GPT-4.182.7

Interactive version: theaggregate.ai/benchmark?slug=prollm-entity-extraction · How the rankings work · Data refreshed daily, snapshot 2026-07-22.