ObGynLongBench - History-Level EHR: leaderboard

Metric: Accuracy (%; full pre-decision EHR history). Source: arxiv.org. 17 models tracked.

Top models

#ModelScore
1Gemini 3 Flash68.7
2GPT-5.4 Mini64.9
3DeepSeek V4 Flash64.4
4Qwen 3.5 35B A3B (Non-reasoning)60.1
5Claude Haiku 4.559.8
6Qwen 3 VL 4B59.8
7Lingshu-7B59.6
8Qwen 3.5 4B (Non-reasoning)59.4
9Qwen 3.5 9B (Non-reasoning)59.3
10Gemma 4 E4B (Non-reasoning)52
11Phi-4 Mini Instruct51.1
12MedGemma 1.5 4B41.3

Interactive version: theaggregate.ai/benchmark?slug=obgynlongbench-history-level-ehr · How It Works · Data refreshed daily, snapshot 2026-09-19.