One Word, Different Action - Decision Sensitivity: leaderboard
Metric: Decision sensitivity (%; exact match of the executable decision (object, object set or ordered action sequence) against deterministic ground truth over 85 physical decision anchors collected with a real robot; share of task-changing instruction pairs where both decisions are correct and differ as the changed constraint requires; temperature 0, at most 512 new tokens, no chain of thought requested; an unmappable output counts as wrong). Source: arxiv.org. Saturation forecast: Around December 2026. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GLM-5 FP8 | 84.3 |
| 2 | Mistral 7B Instruct (v0.3) | 69.6 |
| 3 | OLMo-2-1124-7B-Instruct | 66.7 |
| 4 | Qwen 2.5 7B Instruct | 47.1 |
Interactive version: theaggregate.ai/benchmark?slug=one-word-different-action-decision-sensitivity · How It Works · Data refreshed daily, snapshot 2026-09-26.