LDU-Bench - Cause Analysis: leaderboard
Metric: Rubric score (%; 0.7 x cause-label match, 1 for an exact or alias match and 0.5 within the same supergroup, plus 0.3 x keyword F1 of the short rationale, on 532 expert-reviewed samples; IC-SEM lithography and integrated-circuit review images; one zero-shot run per model at temperature 0 where exposed; invalid outputs stay in the denominator; deterministic scoring, no LLM judge). Source: arxiv.org. Saturation forecast: Around 2028. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.4 | 40.9 |
| 2 | MiniMax-M3 | 39.5 |
| 3 | Qwen 3.6 Plus | 37.6 |
| 4 | GLM-5V Turbo | 29.7 |
| 5 | Claude Opus 4.6 | 25.5 |
Interactive version: theaggregate.ai/benchmark?slug=ldu-bench-cause-analysis · How It Works · Data refreshed daily, snapshot 2026-09-26.