One Word, Different Action - Compositional Accuracy: leaderboard
Metric: Executable-decision accuracy (%; exact match of the executable decision (object, object set or ordered action sequence) against deterministic ground truth over 85 physical decision anchors collected with a real robot; 408 compositional instances in 102 families that join two or three constraints (exclusion, inhibition, cardinality); temperature 0, at most 512 new tokens, no chain of thought requested; an unmappable output counts as wrong). Source: arxiv.org. Saturation forecast: Estimated already saturated. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GLM-5 FP8 | 94.9 |
| 2 | Mistral 7B Instruct (v0.3) | 83.6 |
| 3 | OLMo-2-1124-7B-Instruct | 82.8 |
| 4 | Qwen 2.5 7B Instruct | 64.5 |
Interactive version: theaggregate.ai/benchmark?slug=one-word-different-action-compositional-accuracy · How It Works · Data refreshed daily, snapshot 2026-09-26.