LITMUS - Execution Hallucination: leaderboard
Metric: Execution hallucination rate (%): share of the 117 seed cases whose verbal response and physical OS outcome disagree (claimed but not done, or refused but done), OpenClaw 4.2.0 agent on Ubuntu 24.04 given seed behavior-jailbreak instructions, physical OS state checked by a GPT-4o verifier and analyzer, mean of three runs per case; lower is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen 3.6 Plus | 7.98 |
| 2 | Claude Sonnet 4.6 | 8.07 |
| 3 | Gemini 3.1 Pro (Preview) | 8.89 |
| 4 | DeepSeek V4 Pro | 9.12 |
| 5 | DeepSeek V3.2 | 9.69 |
| 6 | GPT-5.3 Codex | 9.97 |
Interactive version: theaggregate.ai/benchmark?slug=litmus-execution-hallucination · How It Works · Data refreshed daily, snapshot 2026-10-07.