SAD - Predict Words (Situating Prompt): leaderboard

Metric: Self-prediction score (%, higher is better). Source: github.com. 23 models tracked.

Top models

#ModelScore
1Llama 2 7B57.89
2Claude 3.5 Sonnet (20240620)49.64
3Llama 2 13B47.44
4GPT-4 (0613)46.19
5GPT-4 Preview (0125)40.91
6Llama 3 70B Chat38.56
7Claude 3 Opus (20240229)38.41
8GPT-4 Base34.35
9Claude 3 Sonnet (20240229)28.4
10GPT-4o25.72
11Claude Instant 1.224.17
12Claude 3 Haiku (20240307)23.75
13GPT-3.5 Turbo (0613)17.4
14Claude 2.116.26
15davinci-00214.2

Interactive version: theaggregate.ai/benchmark?slug=sad-predict-words-situating-prompt · How It Works · Data refreshed daily, snapshot 2026-09-19.