MedMeta - Title Only (Chain-of-Thought): leaderboard

Metric: LLM-judge score (0-5; mean of Gemini 2.5 Pro, o4-mini and Qwen3-235B judges rating semantic equivalence of the generated conclusion to the meta-analysis's own, 81 meta-analyses; Parametric-CoT: sub-questions answered from the model's own knowledge with a self-feedback loop). Source: arxiv.org. Saturation forecast: Estimated already saturated. 7 models tracked.

Top models

#ModelScore
1Gemini 2.5 Flash2.9
2O4 Mini2.7
3Qwen 3 8B (Non-reasoning)2.27
4Qwen 3 8B (Thinking)2
5DeepSeek R1 0528 Qwen3 8B1.94
6Gemma 3 27B1.77

Interactive version: theaggregate.ai/benchmark?slug=medmeta-title-only-chain-of-thought · How It Works · Data refreshed daily, snapshot 2026-09-25.