SceneActBench - Articulated: leaderboard

Metric: Task score (0-100): 100 S2O articulated objects: rig a closed, unlabelled mesh to reproduce 32 ordered frames of its motion, scored by mean part error against a 1 m reference; every configuration drives one headless Blender through the same MCP tool interface (scene and object inspection, Python execution, rendering) with task step budgets; each native geometric error or F-score is mapped to a 0 to 100 case score against a fixed reference, invalid outputs score 0, cases are averaged per task. Source: arxiv.org. Saturation forecast: Around August 2027. 11 models tracked.

Top models

#ModelScore
1GPT-5.4 (High)73.8
2Claude Opus 4.6 (High)63.7
3GPT-5.4 (Medium)62.3
4Seed 2.0 Pro (High)59.6
5MiniMax M3 (High)58.4
6Claude Sonnet 5 (High)57.8
7Kimi K2.6 (Thinking)57.3
8Gemini 3.1 Pro (Preview) (High)56.5
9Step 3.7 Flash (High)48.3

Interactive version: theaggregate.ai/benchmark?slug=sceneactbench-articulated · How It Works · Data refreshed daily, snapshot 2026-09-29.