JustEval - Factuality — leaderboard

Metric: Score (1-5). Source: allenai.github.io. 16 models tracked.

Top models

#ModelScore
1GPT-4 (0613)4.9
2GPT-4 (0314)4.9
3tulu-2-dpo-70B4.84
4GPT-3.5 Turbo4.83
5Yi 34B (Chat)4.82
6Llama 2 70B Chat4.61
7Llama 2 70B Chat GPTQ4.54
8tulu-2-dpo-7B4.53
9Yi 6B (Chat)4.4
10Vicuna-7B4.33
11Mistral 7B Instruct4.29
12Llama 2 7B Chat4.26

Interactive version: theaggregate.ai/benchmark?slug=justeval-factuality · How the rankings work · Data refreshed daily, snapshot 2026-07-22.