RABBITS MedQA Original — leaderboard

Metric: Accuracy (%). Source: huggingface.co. 26 models tracked.

Top models

#ModelScore
1Gemini 1.5 Flash97.09
2GPT-3.5 Turbo96.3
3GPT-492.33
4GPT-4o90.21
5Gemini 1.5 Pro88.62
6Claude 3 Opus85.71
7Qwen 2 72B75.4
8Llama 3 70B75.13
9Mixtral-8x22B-v0.171.43
10Yi 1.5 34B64.55
11Mixtral 8x7B (v0.1)62.43
12Llama 3 8B60.85
13c4ai-command-r-plus60.32
14Qwen 2 7B58.99
15Phi-3-medium-4k-instruct58.47

Interactive version: theaggregate.ai/benchmark?slug=rabbits-medqa-original · How the rankings work · Data refreshed daily, snapshot 2026-07-22.