Skip to content

The Aggregate

Unified LLM rankings, daily

Models What's New Trends Benchmarks Skill Maps
How It Works Guesswork Metrics
The Aggregate
Loading data...

JustEval — leaderboard

Metric: Avg Score (1-5). Source: allenai.github.io. 16 models tracked.

Top models

#ModelScore
1Yi 34B (Chat)4.87
2tulu-2-dpo-70B4.82
3GPT-4 (0613)4.8
4GPT-4 (0314)4.79
5GPT-3.5 Turbo4.75
6Llama 2 70B Chat4.72
7Llama 2 70B Chat GPTQ4.67
8tulu-2-dpo-7B4.67
9Yi 6B (Chat)4.58
10Llama 2 7B Chat4.47
11Vicuna-7B4.46
12Mistral 7B Instruct4.44

Interactive version: theaggregate.ai/benchmark?slug=justeval · How the rankings work · Data refreshed daily, snapshot 2026-07-22.

Built by Mikhail Doroshenko — AI researcher, co-author of Humanity’s Last Exam. Independent project; no affiliation with any AI lab.