Do LLMs Build World Models From Text? A Multil — leaderboard

Metric: L3 (viewpoint) (self-reported). Source: benchmarklist.com. 14 models tracked.

Top models

#ModelScore
1GPT-4o32.4
2Gemma 3 27B28.7
3DeepSeek V4 Flash26.4
4Gemini 2.5 Flash21.3
5Qwen 2.5 7B Instruct18.9
6Qwen 2.5 32B Instruct14.7
7qwen3.5-flash14.2
8Gemma 2 9B (IT)13.7
9Llama 3.1 8B Instruct13.3
10Mistral 7B Instruct (v0.3)1
11SmolLM2-1.7B-Instruct0

Interactive version: theaggregate.ai/benchmark?slug=do-llms-build-world-models-from-text-a-multil · How the rankings work · Data refreshed daily, snapshot 2026-07-22.