SAD - Stages (Situating Prompt): leaderboard

Metric: Score (%, higher is better). Source: situational-awareness-dataset.org. 21 models tracked.

Top models

#ModelScore
1O1 Preview (2024-09-12)61.52
2O1 Mini (2024-09-12)52.73
3Claude 3 Opus51.09
4Claude 3.5 Sonnet49.31
5Llama 3 70B Chat48.44
6GPT-4 (0613)48.31
7GPT-4 Preview (0125)46.88
8Claude 3 Sonnet44.19
9GPT-4o44.06
10GPT-4 Base43.31
11Llama 2 70B Chat42.31
12Claude 2.141.5
13GPT-3.5 Turbo (0613)41.39
14Llama 2 13B Chat40.83
15Claude Instant 1.239.94

Interactive version: theaggregate.ai/benchmark?slug=sad-stages-situating-prompt · How It Works · Data refreshed daily, snapshot 2026-09-19.