BioMedHop (Direct Prompting) - Path-based Counting: leaderboard

Metric: Accuracy (%; exact count of the distinct diseases that satisfy a typed metapath, after alias deduplication (CountEx); stratified evaluation subsets of the 10,045 instances; direct input-output prompting with no retrieved evidence (the paper's IO baseline), temperature 0, at most 1,024 output tokens). Source: arxiv.org. Saturation forecast: Around 2031. 9 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro13.5
2GPT-4 Turbo12.5
3Claude 3.7 Sonnet12
4Qwen 3 Next 80B A3B Instruct11
5Claude 3.5 Haiku10.2
6MiMo-V2.5-Pro9.5
7GPT-4o Mini8.5
8Qwen 3 4B6

Interactive version: theaggregate.ai/benchmark?slug=biomedhop-direct-prompting-path-based-counting · How It Works · Data refreshed daily, snapshot 2026-09-26.