One Word, Different Action - Decision Sensitivity: leaderboard

Metric: Decision sensitivity (%; exact match of the executable decision (object, object set or ordered action sequence) against deterministic ground truth over 85 physical decision anchors collected with a real robot; share of task-changing instruction pairs where both decisions are correct and differ as the changed constraint requires; temperature 0, at most 512 new tokens, no chain of thought requested; an unmappable output counts as wrong). Source: arxiv.org. Saturation forecast: Around December 2026. 4 models tracked.

Top models

#ModelScore
1GLM-5 FP884.3
2Mistral 7B Instruct (v0.3)69.6
3OLMo-2-1124-7B-Instruct66.7
4Qwen 2.5 7B Instruct47.1

Interactive version: theaggregate.ai/benchmark?slug=one-word-different-action-decision-sensitivity · How It Works · Data refreshed daily, snapshot 2026-09-26.