FLY-EVAL++ - History-Conditioned Prediction - Total: leaderboard

Metric: M1 total score (%). Source: arxiv.org. 21 models tracked.

Top models

#ModelScore
1GPT-597.31
2Gemini 3 Pro97.23
3DeepSeek V397.14
4DeepSeek V3.2 Exp97.08
5Qwen 2.5 32B97.02
6Qwen 3 235B A22B97.01
7Qwen 3 32B96.98
8Claude 3.7 Sonnet (20250219)96.96
9Qwen3-Next-80B96.94
10GPT-4o96.92
11Llama 3.1 405B96.85
12Claude Sonnet 4.596.84
13O4 Mini96.79
14Kimi K2 (Thinking)96.77
15Gemini 2.5 Pro96.76

Interactive version: theaggregate.ai/benchmark?slug=fly-eval-plus-plus-history-conditioned-prediction-total · How It Works · Data refreshed daily, snapshot 2026-09-19.