HREF — leaderboard

HREF evaluates instruction-following models with human response-guided automatic evaluation across 11 task categories.

Metric: Average HREF Score (%). Source: huggingface.co. Status: saturation imminent. 34 models tracked.

Top models

#ModelScore
1Llama 3.1 70B Instruct48.98
2Mistral Large 2 (Jul)48.39
3Qwen 2.5 72B Instruct46.21
4Mistral-Small-Instruct-240942.87
5Qwen 1.5 110B Chat40.76
6Llama 3.1 8B Instruct38.57
7Yi 1.5 34B Chat35.1
8Qwen 2 72B Instruct33.71
9Llama-3.1-Tulu-3-8B33.54
10Phi-3-medium-4k-instruct30.91
11OLMo-2-1124-7B-Instruct28.49
12Llama 2 70B Chat (HF)23.9
13tulu-2-dpo-70B22.67
14Mistral 7B Instruct (v0.3)22.66
15Llama 2 13B Chat Base19.56

Interactive version: theaggregate.ai/benchmark?slug=href · How the rankings work · Data refreshed daily, snapshot 2026-07-22.