FireWorldBench (Structured Records): leaderboard

Metric: Completion accuracy (%; equal-weight mean of the five physical-capability dimensions P1-P5 on the 8,360-item controlled-simulation subset; structured-observation track (textual sensor traces and physical records); each cell is the mean of the matched choice-format score (Jaccard overlap of selected and gold option sets) and open-report score (exact agreement of required structured slots); deterministic evaluator, no LLM judge; failed or unparseable responses score 0). Source: arxiv.org. Saturation forecast: Around 2028. 15 models tracked.

Top models

#ModelScore
1GPT-5.6 Luna55.34
2DeepSeek V3.249.28
3Gemma 3 27B (IT)48.59
4Kimi K247.32
5Qwen 3.5 Plus (2026-02-15)47.25
6Gemini 2.5 Flash46.2
7Qwen 3 32B45.3
8Claude Haiku 4.544.78
9Qwen 3 235B A22B43.78
10Qwen 3 Next 80B A3B Instruct39.48
11GLM 4.5 Air39.23
12Llama 3.3 70B Instruct38.8
13Qwen 3 8B38.07
14Gemma 3 12B (IT)33.7
15Llama 3.1 8B Instruct31.64

Interactive version: theaggregate.ai/benchmark?slug=fireworldbench-structured-records · How It Works · Data refreshed daily, snapshot 2026-09-26.