FireWorldBench - Cross-Field Coupling Understanding (Structured Records): leaderboard

Metric: Completion accuracy (%; P3, cross-field coupling understanding items of the controlled-simulation subset; structured-observation track (textual sensor traces and physical records); each cell is the mean of the matched choice-format score (Jaccard overlap of selected and gold option sets) and open-report score (exact agreement of required structured slots); deterministic evaluator, no LLM judge; failed or unparseable responses score 0). Source: arxiv.org. Saturation forecast: Around 2028. 15 models tracked.

Top models

#ModelScore
1GPT-5.6 Luna55.49
2Gemma 3 27B (IT)46.56
3DeepSeek V3.245.12
4Kimi K244.53
5Qwen 3 235B A22B43.73
6Qwen 3.5 Plus (2026-02-15)42.76
7Gemini 2.5 Flash42.52
8Qwen 3 32B40.16
9Claude Haiku 4.539.92
10GLM 4.5 Air39.84
11Qwen 3 Next 80B A3B Instruct39.53
12Llama 3.3 70B Instruct37.11
13Qwen 3 8B33.88
14Gemma 3 12B (IT)33.3
15Llama 3.1 8B Instruct28.38

Interactive version: theaggregate.ai/benchmark?slug=fireworldbench-cross-field-coupling-understanding-structured-records · How It Works · Data refreshed daily, snapshot 2026-09-26.