3D-DefectBench - Texture (Silver Labels): leaderboard
Metric: Macro MCC (Matthews correlation from -1 to 1, macro-averaged over the four texture defect types) of the model's binary per-defect verdicts against crowd majority labels on the 549-asset silver holdout; the fixed c004 design: six-view oblique RGB turntable renders of each text-to-3D asset and a rubric-guided checklist prompt. Source: arxiv.org. Saturation forecast: Around 2032. 12 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3.1 Pro (Preview) | 0.16 |
| 2 | GPT-5.4 | 0.15 |
| 3 | Gemini 3.1 Flash Lite | 0.15 |
| 4 | Gemini 2.5 Pro | 0.12 |
| 5 | Claude Opus 4.7 | 0.11 |
| 6 | GPT-4o | 0.1 |
| 7 | GPT-5 Mini | 0.07 |
| 8 | Mistral Small 3.1 | 0.06 |
| 9 | Qwen 3.5 397B A17B | 0.05 |
| 10 | Claude Sonnet 4.6 | 0.03 |
| 11 | Qwen 2.5 VL 7B | 0.02 |
| 12 | Claude Haiku 4.5 | 0.02 |
Interactive version: theaggregate.ai/benchmark?slug=3d-defectbench-texture-silver-labels · How It Works · Data refreshed daily, snapshot 2026-09-29.