SAD - Facts: leaderboard

Metric: Score (%, higher is better). Source: situational-awareness-dataset.org. 21 models tracked.

Top models

#ModelScore
1O1 Mini (2024-09-12)77.48
2O1 Preview (2024-09-12)73.9
3GPT-4o65.77
4Claude 3.5 Sonnet63.89
5Claude 3 Opus62.76
6Claude 3 Sonnet61.47
7Claude 2.160.08
8GPT-4 Preview (0125)58.91
9GPT-4 (0613)56.93
10Llama 3 70B Chat56.57
11Claude Instant 1.253.15
12Claude 3 Haiku52.77
13GPT-3.5 Turbo (0613)51.09
14Llama 2 70B Chat50.54
15Llama 2 13B Chat48.63

Interactive version: theaggregate.ai/benchmark?slug=sad-facts · How It Works · Data refreshed daily, snapshot 2026-09-19.