HiEviDR-Bench - Wiki (RAG): leaderboard

Metric: Overall score (0-100): sum of five 0-20 dimension scores (multimodal report quality, evidence traceability, and progressively gated citation, claim and answer scores) judged by Qwen3-VL-235B-A22B-Instruct against the hierarchical evidence graph, on the 1,000-question Wikipedia subset; MLLM with RAG: one-pass report generation over the top-15 retrieved multimodal evidence items; higher is better. Source: arxiv.org. Saturation forecast: Around 2034. 16 models tracked.

Top models

#ModelScore
1GPT-5.6 Sol42.67
2GPT-541.89
3Gemma 4 31B (IT)41.71
4Grok 4.540.82
5GPT-5 Mini40.18
6InternVL3.5-8B39.9
7Qwen 3.5 27B39.35
8gemma-4-E4B-it39.16
9Qwen 3 VL 8B Instruct38.9
10Qwen 3.5 35B A3B37.72
11Qwen 3.5 9B36.07
12Qwen 3.5 4B36.04

Interactive version: theaggregate.ai/benchmark?slug=hievidr-bench-wiki-rag · How It Works · Data refreshed daily, snapshot 2026-09-29.