3D-DefectBench - Geometry (Silver Labels): leaderboard
Metric: Macro MCC (Matthews correlation from -1 to 1, macro-averaged over the five geometry defect types) of the model's binary per-defect verdicts against crowd majority labels on the 549-asset silver holdout; the fixed c004 design: six-view oblique RGB turntable renders of each text-to-3D asset and a rubric-guided checklist prompt. Source: arxiv.org. Saturation forecast: Around 2030. 12 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.4 | 0.24 |
| 2 | Gemini 3.1 Flash Lite | 0.23 |
| 3 | GPT-5 Mini | 0.21 |
| 4 | Gemini 2.5 Pro | 0.21 |
| 5 | Claude Opus 4.7 | 0.21 |
| 6 | Gemini 3.1 Pro (Preview) | 0.21 |
| 7 | Claude Sonnet 4.6 | 0.18 |
| 8 | GPT-4o | 0.17 |
| 9 | Qwen 3.5 397B A17B | 0.15 |
| 10 | Claude Haiku 4.5 | 0.12 |
| 11 | Mistral Small 3.1 | 0.09 |
| 12 | Qwen 2.5 VL 7B | 0.03 |
Interactive version: theaggregate.ai/benchmark?slug=3d-defectbench-geometry-silver-labels · How It Works · Data refreshed daily, snapshot 2026-09-29.