SAD - Introspection (Situating Prompt): leaderboard

Metric: Score (%, higher is better). Source: situational-awareness-dataset.org. 21 models tracked.

Top models

#ModelScore
1Claude 3.5 Sonnet39.38
2GPT-4 (0613)37.75
3Llama 2 7B37.32
4Llama 3 70B Chat35.81
5GPT-4 Base35.04
6GPT-4 Preview (0125)34.93
7Claude 3 Opus34.73
8Llama 2 13B34.68
9Claude Instant 1.230.37
10Claude 3 Sonnet29.97
11GPT-4o29.42
12Claude 3 Haiku28.4
13O1 Mini (2024-09-12)27.86
14Claude 2.125.39
15GPT-3.5 Turbo (0613)24.89

Interactive version: theaggregate.ai/benchmark?slug=sad-introspection-situating-prompt · How It Works · Data refreshed daily, snapshot 2026-09-19.