FireWorldBench - Real-World-Aligned (Structured Records): leaderboard
Metric: Completion accuracy (%; equal-weight mean of P1-P5 on the 714-item real-world-aligned subset, 26 event groups reconstructed from recorded indoor fire experiments; structured-observation track (textual sensor traces and physical records); each cell is the mean of the matched choice-format score (Jaccard overlap of selected and gold option sets) and open-report score (exact agreement of required structured slots); deterministic evaluator, no LLM judge; failed or unparseable responses score 0). Source: arxiv.org. Saturation forecast: Around December 2027. 15 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.6 Luna | 60.16 |
| 2 | Qwen 3.5 Plus (2026-02-15) | 55.35 |
| 3 | Claude Haiku 4.5 | 55.05 |
| 4 | Gemini 2.5 Flash | 54.35 |
| 5 | Llama 3.3 70B Instruct | 51.2 |
| 6 | GLM 4.5 Air | 50.55 |
| 7 | Kimi K2 | 49.54 |
| 8 | DeepSeek V3.2 | 49.24 |
| 9 | Qwen 3 Next 80B A3B Instruct | 48.98 |
| 10 | Qwen 3 32B | 47.64 |
| 11 | Gemma 3 27B (IT) | 44.83 |
| 12 | Qwen 3 235B A22B | 44.47 |
| 13 | Gemma 3 12B (IT) | 41.25 |
| 14 | Qwen 3 8B | 39.09 |
| 15 | Llama 3.1 8B Instruct | 32.54 |
Interactive version: theaggregate.ai/benchmark?slug=fireworldbench-real-world-aligned-structured-records · How It Works · Data refreshed daily, snapshot 2026-09-26.