LLM Self-Modeling - Flip-Decision: leaderboard
Metric: Skill above dummy predictor (-1 to 1, higher is better; 4096-token thinking budget where enabled). Source: arxiv.org. 3 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-OSS-20B (High) | 0.08 |
| 2 | Qwen 3 8B (Thinking) | -0.07 |
| 3 | Llama 3.1 8B Instruct | -0.11 |
Interactive version: theaggregate.ai/benchmark?slug=llm-self-modeling-flip-decision · How It Works · Data refreshed daily, snapshot 2026-09-19.