HiEviDR-Bench - Wiki (Deep Research): leaderboard

Metric: Overall score (0-100): sum of five 0-20 dimension scores (multimodal report quality, evidence traceability, and progressively gated citation, claim and answer scores) judged by Qwen3-VL-235B-A22B-Instruct against the hierarchical evidence graph, on the 1,000-question Wikipedia subset; MLLM with Deep Research: up to three rounds of question refinement with top-20 retrieval, re-ranked to 25 text chunks and 5 images; higher is better. Source: arxiv.org. Saturation forecast: Around 2034. 16 models tracked.

Top models

#ModelScore
1Grok 4.546.41
2GPT-5.6 Sol45.13
3Gemma 4 31B (IT)43.71
4GPT-543.31
5Qwen 3.5 27B42.44
6GPT-5 Mini42.25
7Qwen 3.5 35B A3B41.94
8Qwen 3.5 4B40.35
9gemma-4-E4B-it40.25
10Qwen 3.5 9B39.88
11Qwen 3 VL 8B Instruct39.12
12InternVL3.5-8B38.07

Interactive version: theaggregate.ai/benchmark?slug=hievidr-bench-wiki-deep-research · How It Works · Data refreshed daily, snapshot 2026-09-29.