ForeSci - Prediction Factuality - Strategic Research Planning: leaderboard

Metric: Prediction factuality (claim-level F1 x 100; answer claims supported by, and hidden post-cutoff validation claims covered by, the answer, judged by DeepSeek-V4 with half credit for partial support; 125 tasks ranking research options for a near-term plan; native LLM without retrieval, web search disabled). Source: arxiv.org. Saturation forecast: Around April 2027. 4 models tracked.

Top models

#ModelScore
1Qwen 3 235B A22B46.5
2GPT-5.244.1
3GLM-4.635.9

Interactive version: theaggregate.ai/benchmark?slug=foresci-prediction-factuality-strategic-research-planning · How It Works · Data refreshed daily, snapshot 2026-09-26.