TextClass Benchmark: leaderboard

TextClass Benchmark evaluates LLMs and transformers for social-science text classification across multiple domains and languages, reporting domain-specific Elo leaderboards and a weighted Meta-Elo aggregate.

Metric: Meta-Elo (self-reported). Source: benchmarklist.com. Status: saturated. 108 models tracked.

Top models

#ModelScore
1GPT-4o (2024-05-13)1825.22
2Gemini 1.5 Pro1782.7
3GPT-4 Turbo1781.47
4O1 (2024-12-17)1768.81
5GPT-4.51767.86
6Grok 2 (1212)1758.36
7Llama 3.1 405B1755.81
8GPT-4 (0613)1747.59
9Llama 3.3 70B1746.41
10Grok Beta1741.94
11DeepSeek V31732.54
12Mistral Large 2 (Nov) Instruct (2411)1720.36
13DeepSeek R11718.73
14Gemini 2.0 Flash1701.95
15pixtral-large-24111697.33

Interactive version: theaggregate.ai/benchmark?slug=textclass-benchmark · How It Works · Data refreshed daily, snapshot 2026-09-05.