HiEviDR-Bench - Arxiv (Deep Research): leaderboard

Metric: Overall score (0-100): sum of five 0-20 dimension scores (multimodal report quality, evidence traceability, and progressively gated citation, claim and answer scores) judged by Qwen3-VL-235B-A22B-Instruct against the hierarchical evidence graph, on the 1,000-question arXiv subset; MLLM with Deep Research: up to three rounds of question refinement with top-20 retrieval, re-ranked to 25 text chunks and 5 images; higher is better. Source: arxiv.org. Saturation forecast: Around 2035. 16 models tracked.

Top models

#ModelScore
1Grok 4.535.95
2GPT-5.6 Sol35.13
3Qwen 3.5 9B35.11
4Qwen 3 VL 8B Instruct35.1
5GPT-534.87
6Gemma 4 31B (IT)34.87
7Qwen 3.5 27B34.65
8Qwen 3.5 35B A3B34.61
9Qwen 3.5 4B34.26
10GPT-5 Mini34
11gemma-4-E4B-it33.27
12InternVL3.5-8B33.23

Interactive version: theaggregate.ai/benchmark?slug=hievidr-bench-arxiv-deep-research · How It Works · Data refreshed daily, snapshot 2026-09-29.