MedMeta: leaderboard

Metric: LLM-judge score (0-5; mean of Gemini 2.5 Pro, o4-mini and Qwen3-235B judges rating semantic equivalence of the generated conclusion to the meta-analysis's own, 81 meta-analyses; Golden-RAG: the title and all primary-study abstracts). Source: arxiv.org. Saturation forecast: Estimated already saturated. 7 models tracked.

Top models

#ModelScore
1Gemini 2.5 Flash3.16
2O4 Mini2.79
3Qwen 3 8B (Non-reasoning)2.72
4Gemma 3 27B2.58
5Qwen 3 8B (Thinking)2.56
6DeepSeek R1 0528 Qwen3 8B2.55

Interactive version: theaggregate.ai/benchmark?slug=medmeta · How It Works · Data refreshed daily, snapshot 2026-09-25.