SAD - Self-Recognition (Situating Prompt): leaderboard

Metric: Score (%, higher is better). Source: situational-awareness-dataset.org. 21 models tracked.

Top models

#ModelScore
1O1 Preview (2024-09-12)93
2O1 Mini (2024-09-12)76.13
3Llama 3 70B Chat73.69
4GPT-4 Preview (0125)70.44
5GPT-4o68.63
6Claude 3.5 Sonnet68.13
7Claude 3 Opus63.28
8Claude 2.159.69
9Llama 2 13B Chat56.88
10Claude Instant 1.255.75
11GPT-4 (0613)54
12Llama 2 70B Chat52.88
13Claude 3 Haiku51.69
14GPT-3.5 Turbo (0613)50.81
15davinci-00250.63

Interactive version: theaggregate.ai/benchmark?slug=sad-self-recognition-situating-prompt · How It Works · Data refreshed daily, snapshot 2026-09-19.