JustEval - Helpfulness — leaderboard

Metric: Score (1-5). Source: allenai.github.io. 16 models tracked.

Top models

#ModelScore
1GPT-4 (0314)4.9
2GPT-4 (0613)4.86
3Yi 34B (Chat)4.86
4tulu-2-dpo-70B4.85
5GPT-3.5 Turbo4.81
6tulu-2-dpo-7B4.64
7Llama 2 70B Chat4.58
8Yi 6B (Chat)4.57
9Llama 2 70B Chat GPTQ4.5
10Vicuna-7B4.43
11Mistral 7B Instruct4.36
12Llama 2 7B Chat4.1

Interactive version: theaggregate.ai/benchmark?slug=justeval-helpfulness · How the rankings work · Data refreshed daily, snapshot 2026-07-22.