BioMedHop (Direct Prompting) - Intersection Reasoning: leaderboard

Metric: Accuracy (%; mean of multiple-choice accuracy and alias-normalised open-answer accuracy on Intersection Reasoning (the disease connected to three typed clues); stratified evaluation subsets of the 10,045 instances; direct input-output prompting with no retrieved evidence (the paper's IO baseline), temperature 0, at most 1,024 output tokens). Source: arxiv.org. Saturation forecast: Around 2031. 9 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro22.2
2GPT-4 Turbo21.6
3Claude 3.7 Sonnet21.1
4Qwen 3 Next 80B A3B Instruct19.1
5Claude 3.5 Haiku17.7
6GPT-4o Mini16.7
7MiMo-V2.5-Pro16.2
8Qwen 3 4B9.5

Interactive version: theaggregate.ai/benchmark?slug=biomedhop-direct-prompting-intersection-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-26.