OGCaReBench (RAG Top-1): leaderboard

Metric: Accuracy (%) with the top 1 case report retrieved by BGE-large-en-v1.5 from the 53,617-report corpus placed in the context, on OGCaReBench (639 physician-verified free-text questions on rare, off-guideline cases from PubMed Central case reports, asking for the next clinical step after the standard options are exhausted), answers judged equivalent or mismatched to the case report's step by a GPT-5.2 judge; higher is better. Source: arxiv.org. Saturation forecast: Around April 2027. 9 models tracked.

Top models

#ModelScore
1O4 Mini73.2
2Gemini 2.5 Pro71.8
3Claude Sonnet 4.571
4Claude Sonnet 4 (Thinking)71
5Llama 3.3 70B Instruct69.2
6Llama3-Med42-70B63.4

Interactive version: theaggregate.ai/benchmark?slug=ogcarebench-rag-top-1 · How It Works · Data refreshed daily, snapshot 2026-10-07.