OGCaReBench (RAG Top-3): leaderboard
Metric: Accuracy (%) with the top 3 case reports retrieved by BGE-large-en-v1.5 from the 53,617-report corpus placed in the context, on OGCaReBench (639 physician-verified free-text questions on rare, off-guideline cases from PubMed Central case reports, asking for the next clinical step after the standard options are exhausted), answers judged equivalent or mismatched to the case report's step by a GPT-5.2 judge; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | O4 Mini | 77.8 |
| 2 | Claude Sonnet 4 (Thinking) | 76.1 |
| 3 | Gemini 2.5 Pro | 74.3 |
| 4 | Claude Sonnet 4.5 | 73.7 |
| 5 | Llama 3.3 70B Instruct | 69.5 |
| 6 | Llama3-Med42-70B | 59.2 |
Interactive version: theaggregate.ai/benchmark?slug=ogcarebench-rag-top-3 · How It Works · Data refreshed daily, snapshot 2026-10-07.