TextClass Benchmark — leaderboard

TextClass Benchmark evaluates LLMs and transformers for social-science text classification across multiple domains and languages, reporting domain-specific Elo leaderboards and a weighted Meta-Elo aggregate.

Metric: Meta-Elo (self-reported). Source: benchmarklist.com. Status: saturation imminent. 108 models tracked.

Top models

#ModelScore
1GPT-4o1825.22
2Gemini 1.5 Pro1782.7
3GPT-4 Turbo1781.47
4O11768.81
5GPT-4.51767.86
6Grok 2 (1212)1758.36
7Llama 3.1 405B1755.81
8GPT-41747.59
9Llama 3.3 70B1746.41
10Grok Beta1741.94
11DeepSeek V31732.54
12Mistral Large1720.36
13Gemini 2.0 Flash1701.95
14Gemini 2.0 Flash Lite1687.68
15O3 Mini1684.99

Interactive version: theaggregate.ai/benchmark?slug=textclass-benchmark · How the rankings work · Data refreshed daily, snapshot 2026-07-22.