BehaviorBench - Scientific Workflow Prediction: leaderboard

Metric: BLEURT score of generated research-workflow aspects (key idea, method, outcome, projected impact and title, each from the preceding aspects) against 2025 American Economic Review and Nature Human Behaviour papers; higher is better. Source: arxiv.org. Saturation forecast: Around August 2027. 18 models tracked.

Top models

#ModelScore
1Claude Opus 4.60.48
2Claude Sonnet 4.60.47
3Gemini 3.1 Pro (Preview)0.47
4GPT-5.4 (High)0.46
5GPT-4.10.46
6GPT-5.4 Mini (High)0.45
7Qwen 3 4B0.45
8Gemini 3.1 Flash Lite (Preview)0.43
9Llama 3.3 70B Instruct0.43
10DeepSeek V3.20.43
11Claude Haiku 4.50.43

Interactive version: theaggregate.ai/benchmark?slug=behaviorbench-scientific-workflow-prediction · How It Works · Data refreshed daily, snapshot 2026-09-29.