PaperMind - Experimental Interpretation: leaderboard

Metric: Judge score for writing the analysis paragraph for an experimental figure or table given the paper's introduction, the paper's own discussion paragraph as reference; rated by a GPT-4o judge on a 5-point scale (1-5) against the reference, averaged over the seven scientific domains; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 7 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro2.25
2Claude 3.5 Sonnet2.13
3GPT-4o Mini1.92
4Qwen 3 VL 4B Instruct1.9
5Claude 3 Haiku1.82
6Gemma 3 4B (IT)1.73

Interactive version: theaggregate.ai/benchmark?slug=papermind-experimental-interpretation · How It Works · Data refreshed daily, snapshot 2026-10-07.