HyperGVL - Reasoning: leaderboard
Metric: Average accuracy (%) over the six reasoning tasks (order-weighted shortest path and maximum flow, isomorphism, 3-coloring, strict hypercycle, Hamilton path), 7,000 questions per task (200 problems on synthetic and real-world hypergraphs, each posed in 35 textual, visual and combined hypergraph representations), accuracy (%) averaged over the representations, temperature 0.8, top-p 0.95; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 12 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Flash | 62.04 |
| 2 | GPT-4o | 35 |
| 3 | Gemma 3 12B | 25.53 |
| 4 | Qwen 2.5 VL 7B Instruct | 21.3 |
| 5 | InternVL3-8B | 20.52 |
Interactive version: theaggregate.ai/benchmark?slug=hypergvl-reasoning · How It Works · Data refreshed daily, snapshot 2026-10-07.