FireWorldBench - Temporal Evolution Forecasting (Structured Records): leaderboard

Metric: Completion accuracy (%; P1, temporal evolution forecasting items of the controlled-simulation subset; structured-observation track (textual sensor traces and physical records); each cell is the mean of the matched choice-format score (Jaccard overlap of selected and gold option sets) and open-report score (exact agreement of required structured slots); deterministic evaluator, no LLM judge; failed or unparseable responses score 0). Source: arxiv.org. Saturation forecast: Around December 2026. 15 models tracked.

Top models

#ModelScore
1GPT-5.6 Luna66.59
2Kimi K261.28
3Qwen 3.5 Plus (2026-02-15)60.77
4DeepSeek V3.260.7
5Gemini 2.5 Flash59.92
6Gemma 3 27B (IT)59.04
7Qwen 3 32B58.3
8Claude Haiku 4.556.56
9Qwen 3 235B A22B55.38
10GLM 4.5 Air49.53
11Qwen 3 Next 80B A3B Instruct48.41
12Llama 3.3 70B Instruct47.08
13Qwen 3 8B46.55
14Gemma 3 12B (IT)34.67
15Llama 3.1 8B Instruct31.93

Interactive version: theaggregate.ai/benchmark?slug=fireworldbench-temporal-evolution-forecasting-structured-records · How It Works · Data refreshed daily, snapshot 2026-09-26.