BioMedHop (Direct Prompting): leaderboard

Metric: Overall accuracy (%; mean of the four task-family scores, which are MCQ/open-answer accuracy averages and exact-count accuracy; stratified evaluation subsets of the 10,045 instances; direct input-output prompting with no retrieved evidence (the paper's IO baseline), temperature 0, at most 1,024 output tokens). Source: arxiv.org. Saturation forecast: Around 2032. 9 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro18.6
2GPT-4 Turbo18.2
3Claude 3.7 Sonnet17.8
4Qwen 3 Next 80B A3B Instruct16.1
5Claude 3.5 Haiku14.5
6MiMo-V2.5-Pro13.7
7GPT-4o Mini13.2
8Qwen 3 4B9.3

Interactive version: theaggregate.ai/benchmark?slug=biomedhop-direct-prompting · How It Works · Data refreshed daily, snapshot 2026-09-26.