3D-DefectBench - Texture (Expert Labels): leaderboard
Metric: Macro MCC (Matthews correlation from -1 to 1, macro-averaged over the four texture defect types) of the model's binary per-defect verdicts against expert labels on the cells both experts agree on (129-asset expert split); the fixed c004 design: six-view oblique RGB turntable renders of each text-to-3D asset and a rubric-guided checklist prompt. Source: arxiv.org. Saturation forecast: Around September 2028. 12 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3.1 Flash Lite | 0.41 |
| 2 | Gemini 3.1 Pro (Preview) | 0.39 |
| 3 | Claude Opus 4.7 | 0.3 |
| 4 | Gemini 2.5 Pro | 0.28 |
| 5 | GPT-4o | 0.26 |
| 6 | Claude Sonnet 4.6 | 0.23 |
| 7 | GPT-5.4 | 0.22 |
| 8 | Qwen 3.5 397B A17B | 0.14 |
| 9 | GPT-5 Mini | 0.12 |
| 10 | Qwen 2.5 VL 7B | 0.08 |
| 11 | Mistral Small 3.1 | 0.07 |
| 12 | Claude Haiku 4.5 | 0.03 |
Interactive version: theaggregate.ai/benchmark?slug=3d-defectbench-texture-expert-labels · How It Works · Data refreshed daily, snapshot 2026-09-29.