BioMedHop (Direct Prompting) - Entity Pair Matching: leaderboard

Metric: Accuracy (%; mean of multiple-choice accuracy and alias-normalised open-answer accuracy on Entity Pair Matching (the shared biomedical neighbour of two anchor entities); stratified evaluation subsets of the 10,045 instances; direct input-output prompting with no retrieved evidence (the paper's IO baseline), temperature 0, at most 1,024 output tokens). Source: arxiv.org. Saturation forecast: Around 2031. 9 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro13.3
2GPT-4 Turbo12.7
3Claude 3.7 Sonnet12.4
4Qwen 3 Next 80B A3B Instruct10.8
5MiMo-V2.5-Pro9.4
6Claude 3.5 Haiku8.9
7GPT-4o Mini7.8
8Qwen 3 4B5.9

Interactive version: theaggregate.ai/benchmark?slug=biomedhop-direct-prompting-entity-pair-matching · How It Works · Data refreshed daily, snapshot 2026-09-26.