LDU-Bench - Cause Analysis: leaderboard

Metric: Rubric score (%; 0.7 x cause-label match, 1 for an exact or alias match and 0.5 within the same supergroup, plus 0.3 x keyword F1 of the short rationale, on 532 expert-reviewed samples; IC-SEM lithography and integrated-circuit review images; one zero-shot run per model at temperature 0 where exposed; invalid outputs stay in the denominator; deterministic scoring, no LLM judge). Source: arxiv.org. Saturation forecast: Around 2028. 6 models tracked.

Top models

#ModelScore
1GPT-5.440.9
2MiniMax-M339.5
3Qwen 3.6 Plus37.6
4GLM-5V Turbo29.7
5Claude Opus 4.625.5

Interactive version: theaggregate.ai/benchmark?slug=ldu-bench-cause-analysis · How It Works · Data refreshed daily, snapshot 2026-09-26.