MUSE (Text-to-CAD) - Geometric Validity: leaderboard

Metric: Geometric validity (%): share of the 106 design specifications whose exported STEP model passes all four binary geometry checks (watertight, manifold, self-intersection free, overlap free), one CadQuery script per design specification, executed in a sandbox; funnel evaluation, so a sample that fails a stage scores zero on every later stage; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 15 models tracked.

Top models

#ModelScore
1GPT-5.568.87
2Gemini 3.1 Pro (Preview)58.49
3Claude Opus 4.758.49
4GLM-5.127.36
5Claude 3.7 Sonnet23.58
6Llama 3.1 70B23.58
7Qwen 2.5 72B20.75
8GPT-4o13.21
9Qwen 3.5 122B A10B13.21
10MiniMax-M2.710.38
11MiniMax-M2.510.38
12Qwen 3.6 35B A3B5.66
13GLM-4.7 Flash5.66
14Llama 3.1 8B1.89

Interactive version: theaggregate.ai/benchmark?slug=muse-text-to-cad-geometric-validity · How It Works · Data refreshed daily, snapshot 2026-10-07.