Open LLM Leaderboard v1 - HellaSwag: leaderboard
Metric: Normalized accuracy (%) (10-shot). Source: huggingface.co. Saturation forecast: Estimated already saturated. 7252 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | LLaMAntino-3-ANITA-8B-Inst-DPO-ITA | 92.75 |
| 2 | UNA-ThePitbull-21.4B-v2 | 91.79 |
| 3 | free-evo-qwen72B-v0.8-re | 91.34 |
| 4 | Rhea-72B-v0.5 | 91.15 |
| 5 | luxia-21.4B-alignment-v1.2 | 90.86 |
| 6 | final_model_test_v2 | 90.86 |
| 7 | Le_Triomphant-ECE-TW3 | 90.3 |
| 8 | test1 | 89.52 |
| 9 | DARE_TIES_13B | 89.5 |
| 10 | Llama-3-SauerkrautLM-8B-Instruct | 89.41 |
| 11 | MoE_13B_DPO | 89.39 |
| 12 | testmerge-7B | 89.37 |
| 13 | CarbonBeagle-11B-truthy | 89.31 |
| 14 | Truthful_DPO_TomGrc_FusionNet_7Bx2_MoE_13B | 89.3 |
| 15 | tulu-2-dpo-70B-ExPO | 89.29 |
Interactive version: theaggregate.ai/benchmark?slug=open-llm-leaderboard-v1-hellaswag · How It Works · Data refreshed daily, snapshot 2026-09-23.