Skip to content

The Aggregate

Unified LLM rankings, daily

Models What's New Trends Benchmarks Skill Maps
How It Works Guesswork Metrics
The Aggregate
Loading data...

JustEval - Safety — leaderboard

Metric: Score (1-5). Source: allenai.github.io. 16 models tracked.

Top models

#ModelScore
1Llama 2 7B Chat5
2Llama 2 70B Chat5
3Llama 2 70B Chat GPTQ5
4tulu-2-dpo-70B4.99
5GPT-4 (0613)4.97
6GPT-3.5 Turbo4.94
7Yi 34B (Chat)4.92
8tulu-2-dpo-7B4.88
9Mistral 7B Instruct4.75
10GPT-4 (0314)4.74
11Yi 6B (Chat)4.67
12Vicuna-7B4.6

Interactive version: theaggregate.ai/benchmark?slug=justeval-safety · How the rankings work · Data refreshed daily, snapshot 2026-07-22.

Built by Mikhail Doroshenko — AI researcher, co-author of Humanity’s Last Exam. Independent project; no affiliation with any AI lab.