Ishigaki-IDS-Bench: leaderboard

Metric: Facet F1 (%): F1 of exact matches between generated and gold IDS facet blocks (IFC version, entity, attribute and property facets with their applicability or requirement role and constraints), macro-averaged over examples, an output with no extractable IDS counting as invalid, on Ishigaki-IDS-Bench, 166 expert-authored examples (83 practical BIM information-requirement scenarios in Japanese and English, CSV or natural-language input, single- or multi-turn) whose output must be an IDS 1.0 XML file matching a gold IDS, zero-shot; higher is better. Source: arxiv.org. Saturation forecast: Around May 2027. 10 models tracked.

Top models

#ModelScore
1GPT-5.5 (xHigh)65.6
2Gemini 3.1 Pro (Preview)52.2
3Kimi K2.648.4
4Claude Opus 4.542.2
5DeepSeek V4 Pro40
6Qwen 3.5 397B A17B26.8
7Qwen 3 14B21.3

Interactive version: theaggregate.ai/benchmark?slug=ishigaki-ids-bench · How It Works · Data refreshed daily, snapshot 2026-10-07.