HiEviDR-Bench (Deep Research): leaderboard

Metric: Overall score (0-100): sum of five 0-20 dimension scores (multimodal report quality, evidence traceability, and progressively gated citation, claim and answer scores) judged by Qwen3-VL-235B-A22B-Instruct against the hierarchical evidence graph, averaged over the Wikipedia and arXiv subsets; MLLM with Deep Research: up to three rounds of question refinement with top-20 retrieval, re-ranked to 25 text chunks and 5 images; higher is better. Source: arxiv.org. Saturation forecast: Around 2034. 16 models tracked.

Top models

#ModelScore
1Grok 4.541.18
2GPT-5.6 Sol40.13
3Gemma 4 31B (IT)39.29
4GPT-539.09
5Qwen 3.5 27B38.54
6Qwen 3.5 35B A3B38.28
7GPT-5 Mini38.13
8Qwen 3.5 9B37.5
9Qwen 3.5 4B37.31
10Qwen 3 VL 8B Instruct37.11
11gemma-4-E4B-it36.76
12InternVL3.5-8B35.65

Interactive version: theaggregate.ai/benchmark?slug=hievidr-bench-deep-research · How It Works · Data refreshed daily, snapshot 2026-09-29.