One Word, Different Action - Atomic Accuracy: leaderboard
Metric: Executable-decision accuracy (%; exact match of the executable decision (object, object set or ordered action sequence) against deterministic ground truth over 85 physical decision anchors collected with a real robot; 3,952 atomic instances in 988 task families, one constraint each (inhibition, exclusion, cardinality or temporal ordering); temperature 0, at most 512 new tokens, no chain of thought requested; an unmappable output counts as wrong). Source: arxiv.org. Saturation forecast: Estimated already saturated. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GLM-5 FP8 | 99.4 |
| 2 | Mistral 7B Instruct (v0.3) | 97.8 |
| 3 | OLMo-2-1124-7B-Instruct | 87.8 |
| 4 | Qwen 2.5 7B Instruct | 84.4 |
Interactive version: theaggregate.ai/benchmark?slug=one-word-different-action-atomic-accuracy · How It Works · Data refreshed daily, snapshot 2026-09-26.