LDU-Bench - Defect Triage: leaderboard

Metric: Macro-F1 (%; binary good versus defect decision on 1,761 images; IC-SEM lithography and integrated-circuit review images; one zero-shot run per model at temperature 0 where exposed; invalid outputs stay in the denominator; deterministic scoring, no LLM judge). Source: arxiv.org. Saturation forecast: Estimated already saturated. 6 models tracked.

Top models

#ModelScore
1GPT-5.493.2
2MiniMax-M390.9
3GLM-5V Turbo88.4
4Claude Opus 4.687.6
5Qwen 3.6 Plus85.6

Interactive version: theaggregate.ai/benchmark?slug=ldu-bench-defect-triage · How It Works · Data refreshed daily, snapshot 2026-09-26.