ForeSci - Reviewer Persuasiveness: leaderboard
Metric: Reviewer persuasiveness (x 100; holistic rubric score from a DeepSeek-V4 virtual reviewer given the task, the pre-cutoff evidence and the hidden targets, mean of repeated runs and of the four task families, 500 tasks; native LLM without retrieval, web search disabled). Source: arxiv.org. Saturation forecast: Around March 2027. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.2 | 84.6 |
| 2 | Qwen 3 235B A22B | 78.6 |
| 3 | GLM-4.6 | 67.4 |
Interactive version: theaggregate.ai/benchmark?slug=foresci-reviewer-persuasiveness · How It Works · Data refreshed daily, snapshot 2026-09-26.