HiEviDR-Bench - Arxiv (RAG): leaderboard

Metric: Overall score (0-100): sum of five 0-20 dimension scores (multimodal report quality, evidence traceability, and progressively gated citation, claim and answer scores) judged by Qwen3-VL-235B-A22B-Instruct against the hierarchical evidence graph, on the 1,000-question arXiv subset; MLLM with RAG: one-pass report generation over the top-15 retrieved multimodal evidence items; higher is better. Source: arxiv.org. Saturation forecast: Around 2035. 16 models tracked.

Top models

#ModelScore
1GPT-5.6 Sol34.03
2Gemma 4 31B (IT)33.82
3GPT-533.34
4Qwen 3.5 27B32.83
5Grok 4.532.79
6InternVL3.5-8B32.5
7gemma-4-E4B-it32.49
8Qwen 3 VL 8B Instruct32.18
9Qwen 3.5 4B31.97
10Qwen 3.5 35B A3B30.82
11Qwen 3.5 9B30.19

Interactive version: theaggregate.ai/benchmark?slug=hievidr-bench-arxiv-rag · How It Works · Data refreshed daily, snapshot 2026-09-29.