RoboAbstention - Contradictory Instructions: leaderboard

Metric: Abstention rate (%) on the 837 Contradictory Instructions instructions, image plus instruction to an embodied VLM planner, every instruction warranting abstention, act or abstain judged by GPT-5.4 mini (97.5 percent agreement with human labels), API default sampling; higher is better. Source: arxiv.org. Saturation forecast: Around 2035. 11 models tracked.

Top models

#ModelScore
1Gemini 2.5 Flash36.3
2GPT-5.429
3Llama 4 Maverick28.6
4Claude Sonnet 4.626.5
5GPT-5.4 Nano22.6
6GPT-5.4 Mini (Low)22
7GPT-5.4 Mini (Medium)20.5
8GPT-5.4 Mini (High)17.2
9GPT-5.4 Mini (Non-reasoning)16.8
10Qwen 3.5 27B5.4

Interactive version: theaggregate.ai/benchmark?slug=roboabstention-contradictory-instructions · How It Works · Data refreshed daily, snapshot 2026-10-07.