HyperGVL - Reasoning: leaderboard

Metric: Average accuracy (%) over the six reasoning tasks (order-weighted shortest path and maximum flow, isomorphism, 3-coloring, strict hypercycle, Hamilton path), 7,000 questions per task (200 problems on synthetic and real-world hypergraphs, each posed in 35 textual, visual and combined hypergraph representations), accuracy (%) averaged over the representations, temperature 0.8, top-p 0.95; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 12 models tracked.

Top models

#ModelScore
1Gemini 3 Flash62.04
2GPT-4o35
3Gemma 3 12B25.53
4Qwen 2.5 VL 7B Instruct21.3
5InternVL3-8B20.52

Interactive version: theaggregate.ai/benchmark?slug=hypergvl-reasoning · How It Works · Data refreshed daily, snapshot 2026-10-07.