HiEviDR-Bench (RAG): leaderboard

Metric: Overall score (0-100): sum of five 0-20 dimension scores (multimodal report quality, evidence traceability, and progressively gated citation, claim and answer scores) judged by Qwen3-VL-235B-A22B-Instruct against the hierarchical evidence graph, averaged over the Wikipedia and arXiv subsets; MLLM with RAG: one-pass report generation over the top-15 retrieved multimodal evidence items; higher is better. Source: arxiv.org. Saturation forecast: Around 2035. 16 models tracked.

Top models

#ModelScore
1GPT-5.6 Sol38.35
2Gemma 4 31B (IT)37.76
3GPT-537.62
4Grok 4.536.82
5GPT-5 Mini36.74
6InternVL3.5-8B36.2
7Qwen 3.5 27B36.09
8gemma-4-E4B-it35.83
9Qwen 3 VL 8B Instruct35.54
10Qwen 3.5 35B A3B34.27
11Qwen 3.5 4B34
12Qwen 3.5 9B33.13

Interactive version: theaggregate.ai/benchmark?slug=hievidr-bench-rag · How It Works · Data refreshed daily, snapshot 2026-09-29.