HellaSwag: leaderboard

Commonsense natural language inference benchmark with 70K questions. Given a scenario, models must pick the most plausible continuation. Heavily saturated by modern models.

Metric: Accuracy (%). Source: epoch.ai. Status: saturated. 76 models tracked.

Top models

#ModelScore
1Median Human95.6
2GPT-4 (0314)95.3
3Llama 3.1 405B89.2
4DeepSeek V388.9
5DeepSeek V287.1
6Mixtral 8x7B (v0.1)86.7
7text-davinci-00385.5
8Llama 2 70B Base85.3
9falcon-40B85.28
10Qwen 2.5 72B84.8
11LLaMA-65B84.2
12Qwen 2.5 Coder 32B83
13falcon-11B82.91
14Phi-3-medium-128k-instruct82.4
15gemma-7B82.2

Interactive version: theaggregate.ai/benchmark?slug=hellaswag · How It Works · Data refreshed daily, snapshot 2026-09-05.