JustEval - Depth — leaderboard

Metric: Score (1-5). Source: allenai.github.io. 16 models tracked.

Top models

#ModelScore
1Yi 34B (Chat)4.79
2GPT-4 (0314)4.57
3tulu-2-dpo-70B4.57
4GPT-4 (0613)4.49
5Yi 6B (Chat)4.39
6Llama 2 70B Chat4.38
7tulu-2-dpo-7B4.36
8GPT-3.5 Turbo4.33
9Llama 2 70B Chat GPTQ4.28
10Vicuna-7B4.04
11Llama 2 7B Chat3.91
12Mistral 7B Instruct3.89

Interactive version: theaggregate.ai/benchmark?slug=justeval-depth · How the rankings work · Data refreshed daily, snapshot 2026-07-22.