ForeSci - Prediction Factuality - Direction Forecasting: leaderboard
Metric: Prediction factuality (claim-level F1 x 100; answer claims supported by, and hidden post-cutoff validation claims covered by, the answer, judged by DeepSeek-V4 with half credit for partial support; 125 tasks choosing the candidate direction most likely to gain momentum; native LLM without retrieval, web search disabled). Source: arxiv.org. Saturation forecast: Around March 2027. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.2 | 67 |
| 2 | Qwen 3 235B A22B | 59.8 |
| 3 | GLM-4.6 | 52.3 |
Interactive version: theaggregate.ai/benchmark?slug=foresci-prediction-factuality-direction-forecasting · How It Works · Data refreshed daily, snapshot 2026-09-26.