OGCaReBench (RAG Top-3): leaderboard

Metric: Accuracy (%) with the top 3 case reports retrieved by BGE-large-en-v1.5 from the 53,617-report corpus placed in the context, on OGCaReBench (639 physician-verified free-text questions on rare, off-guideline cases from PubMed Central case reports, asking for the next clinical step after the standard options are exhausted), answers judged equivalent or mismatched to the case report's step by a GPT-5.2 judge; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 9 models tracked.

Top models

#ModelScore
1O4 Mini77.8
2Claude Sonnet 4 (Thinking)76.1
3Gemini 2.5 Pro74.3
4Claude Sonnet 4.573.7
5Llama 3.3 70B Instruct69.5
6Llama3-Med42-70B59.2

Interactive version: theaggregate.ai/benchmark?slug=ogcarebench-rag-top-3 · How It Works · Data refreshed daily, snapshot 2026-10-07.