LDU-Bench - Coarse Localization: leaderboard
Metric: Mean DICU (%; harmonic mean of box IoU and ground-truth mask coverage for one xyxy box per image on 984 mask-annotated images, unparsable boxes score 0; IC-SEM lithography and integrated-circuit review images; one zero-shot run per model at temperature 0 where exposed; invalid outputs stay in the denominator; deterministic scoring, no LLM judge). Source: arxiv.org. Saturation forecast: Around June 2028. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.4 | 49.2 |
| 2 | GLM-5V Turbo | 35.3 |
| 3 | MiniMax-M3 | 33.8 |
| 4 | Qwen 3.6 Plus | 26.9 |
| 5 | Claude Opus 4.6 | 25.4 |
Interactive version: theaggregate.ai/benchmark?slug=ldu-bench-coarse-localization · How It Works · Data refreshed daily, snapshot 2026-09-26.