FireWorldBench - Causal Mechanism Attribution (Structured Records): leaderboard

Metric: Completion accuracy (%; P4, causal mechanism attribution items of the controlled-simulation subset; structured-observation track (textual sensor traces and physical records); each cell is the mean of the matched choice-format score (Jaccard overlap of selected and gold option sets) and open-report score (exact agreement of required structured slots); deterministic evaluator, no LLM judge; failed or unparseable responses score 0). Source: arxiv.org. Saturation forecast: Around 2031. 15 models tracked.

Top models

#ModelScore
1GPT-5.6 Luna52.14
2DeepSeek V3.246.91
3Qwen 3 8B42.98
4Qwen 3 32B42.67
5Claude Haiku 4.542.03
6Kimi K240.98
7Qwen 3.5 Plus (2026-02-15)40.61
8Qwen 3 235B A22B40.56
9Llama 3.3 70B Instruct38.02
10GLM 4.5 Air37.49
11Gemini 2.5 Flash37.32
12Llama 3.1 8B Instruct37.16
13Gemma 3 27B (IT)37.08
14Gemma 3 12B (IT)34.03
15Qwen 3 Next 80B A3B Instruct32.91

Interactive version: theaggregate.ai/benchmark?slug=fireworldbench-causal-mechanism-attribution-structured-records · How It Works · Data refreshed daily, snapshot 2026-09-26.