LLM Self-Modeling: leaderboard

Evaluating and Improving LLM Self-Modeling (2026), Table 16: seven open models predicting their own input-output behavior across math, code, safety and fairness. Strict task scores minus task/model-specific dummy baselines, averaged over five seeds. Nine-task E1-E9 aggregate. Uses unmodified instruction-tuned checkpoints with the published inference settings; no self-modeling fine-tunes or values estimated from plotted bars.

Metric: Skill above dummy predictor (-1 to 1, higher is better; 4096-token thinking budget where enabled). Source: arxiv.org. Status: saturation imminent. 7 models tracked.

Top models

#ModelScore
1DeepSeek V3.1 (Thinking)0.15
2Qwen 3 8B (Thinking)0.11
3Qwen 3.5 35B A3B (Thinking)0.11
4GPT-OSS-20B (High)0.08
5Qwen 3.5 4B (Thinking)0.08
6Llama 3.3 70B Instruct0.06
7Llama 3.1 8B Instruct-0.01

Interactive version: theaggregate.ai/benchmark?slug=llm-self-modeling · How It Works · Data refreshed daily, snapshot 2026-09-19.