MedMeta - Title Only: leaderboard

Metric: LLM-judge score (0-5; mean of Gemini 2.5 Pro, o4-mini and Qwen3-235B judges rating semantic equivalence of the generated conclusion to the meta-analysis's own, 81 meta-analyses; zero-shot from the title alone). Source: arxiv.org. Saturation forecast: Estimated already saturated. 7 models tracked.

Top models

#ModelScore
1Gemini 2.5 Flash2.1
2O4 Mini2
3Qwen 3 8B (Non-reasoning)1.7
4Gemma 3 27B1.6
5Qwen 3 8B (Thinking)1.5
6DeepSeek R1 0528 Qwen3 8B1.3

Interactive version: theaggregate.ai/benchmark?slug=medmeta-title-only · How It Works · Data refreshed daily, snapshot 2026-09-25.