BIM-Edit - Semantics: leaderboard
Metric: Semantic score (0-100): IFC class match and task-relevant property match of the edited elements against the reference (unmatched references score zero), on 324 natural-language create, update and delete tasks on 11 real and 36 synthetic IFC building models (direct, spatial and topological instructions), the model edits the IFC file as an agent that runs Python through a single IfcOpenShell execution tool with a 20-call budget, temperature 0 where settable; mean over tasks; higher is better. Source: arxiv.org. Saturation forecast: Around 2031. 7 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Flash (Preview) | 41.81 |
| 2 | GPT-5.4 Mini | 38.8 |
| 3 | GPT-5.4 Pro (xHigh) | 38.69 |
| 4 | Qwen 3.6 Plus | 35.93 |
| 5 | DeepSeek V3.2 | 34.8 |
| 6 | Claude Sonnet 4.6 | 33.85 |
| 7 | Gemma 4 31B (IT) | 29.94 |
Interactive version: theaggregate.ai/benchmark?slug=bim-edit-semantics · How It Works · Data refreshed daily, snapshot 2026-09-29.