One Word, Different Action - Atomic Accuracy: leaderboard

Metric: Executable-decision accuracy (%; exact match of the executable decision (object, object set or ordered action sequence) against deterministic ground truth over 85 physical decision anchors collected with a real robot; 3,952 atomic instances in 988 task families, one constraint each (inhibition, exclusion, cardinality or temporal ordering); temperature 0, at most 512 new tokens, no chain of thought requested; an unmappable output counts as wrong). Source: arxiv.org. Saturation forecast: Estimated already saturated. 4 models tracked.

Top models

#ModelScore
1GLM-5 FP899.4
2Mistral 7B Instruct (v0.3)97.8
3OLMo-2-1124-7B-Instruct87.8
4Qwen 2.5 7B Instruct84.4

Interactive version: theaggregate.ai/benchmark?slug=one-word-different-action-atomic-accuracy · How It Works · Data refreshed daily, snapshot 2026-09-26.