LDU-Bench - Defect Triage: leaderboard
Metric: Macro-F1 (%; binary good versus defect decision on 1,761 images; IC-SEM lithography and integrated-circuit review images; one zero-shot run per model at temperature 0 where exposed; invalid outputs stay in the denominator; deterministic scoring, no LLM judge). Source: arxiv.org. Saturation forecast: Estimated already saturated. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.4 | 93.2 |
| 2 | MiniMax-M3 | 90.9 |
| 3 | GLM-5V Turbo | 88.4 |
| 4 | Claude Opus 4.6 | 87.6 |
| 5 | Qwen 3.6 Plus | 85.6 |
Interactive version: theaggregate.ai/benchmark?slug=ldu-bench-defect-triage · How It Works · Data refreshed daily, snapshot 2026-09-26.