P3D-Bench - Assembly-3D Parts: leaderboard

Metric: Part bucket score (0-100) (part-match F1 and part shape F-score after a fixed MLLM decomposes each valid assembly into parts matched one-to-one to the ground-truth parts) on 203 annotated multi-part assemblies (CadQuery and OpenSCAD outputs), the model writes a parametric CAD program that is compiled, meshed and aligned to the ground truth; invalid outputs take the worst value; each general model runs with its maximum thinking budget; averaged over the output formats; higher is better. Source: arxiv.org. Saturation forecast: Around May 2027. 8 models tracked.

Top models

#ModelScore
1GPT-5.5 (xHigh)62.9
2Gemini 3.1 Pro (Preview) (High)61.8
3Claude Opus 4.6 (Max)57.3
4MiMo-V2-Omni38.8
5Qwen 3.6 Plus (Thinking)35.1
6GLM-5V Turbo33.8
7Seed 2.0 Pro (High)32.2

Interactive version: theaggregate.ai/benchmark?slug=p3d-bench-assembly-3d-parts · How It Works · Data refreshed daily, snapshot 2026-09-29.