ForeSci - Reviewer Persuasiveness: leaderboard

Metric: Reviewer persuasiveness (x 100; holistic rubric score from a DeepSeek-V4 virtual reviewer given the task, the pre-cutoff evidence and the hidden targets, mean of repeated runs and of the four task families, 500 tasks; native LLM without retrieval, web search disabled). Source: arxiv.org. Saturation forecast: Around March 2027. 4 models tracked.

Top models

#ModelScore
1GPT-5.284.6
2Qwen 3 235B A22B78.6
3GLM-4.667.4

Interactive version: theaggregate.ai/benchmark?slug=foresci-reviewer-persuasiveness · How It Works · Data refreshed daily, snapshot 2026-09-26.