FireWorldBench - Physical Field Perception and Grounding (Structured Records): leaderboard

Metric: Completion accuracy (%; P2, physical field perception and grounding items of the controlled-simulation subset; structured-observation track (textual sensor traces and physical records); each cell is the mean of the matched choice-format score (Jaccard overlap of selected and gold option sets) and open-report score (exact agreement of required structured slots); deterministic evaluator, no LLM judge; failed or unparseable responses score 0). Source: arxiv.org. Saturation forecast: Around June 2028. 15 models tracked.

Top models

#ModelScore
1GPT-5.6 Luna54.44
2Claude Haiku 4.553.95
3Gemma 3 27B (IT)53.4
4Kimi K251.8
5Gemini 2.5 Flash50.81
6DeepSeek V3.249.23
7Qwen 3.5 Plus (2026-02-15)47.2
8Qwen 3 32B45.76
9Llama 3.3 70B Instruct44
10Qwen 3 235B A22B42.76
11Qwen 3 Next 80B A3B Instruct42.61
12GLM 4.5 Air39.76
13Gemma 3 12B (IT)38.57
14Qwen 3 8B35.34
15Llama 3.1 8B Instruct31.35

Interactive version: theaggregate.ai/benchmark?slug=fireworldbench-physical-field-perception-and-grounding-structured-records · How It Works · Data refreshed daily, snapshot 2026-09-26.