CIPHER Cross-Record Inference - MIMIC (Few-Shot): leaderboard

Metric: Harmonic answer score (%; per question the harmonic mean of token overlap with the reference and Gemini 2.5 Flash semantic acceptance, zero when rejected, averaged; MIMIC clinical records with redacted narratives; the full support-controlled context of 39-42 records including distractors; two-example few-shot prompting; greedy decoding). Source: arxiv.org. Saturation forecast: Around 2028. 5 models tracked.

Top models

#ModelScore
1GPT-OSS-20B32.3
2GPT-OSS-120B26.2
3Llama 3.3 70B Instruct25.8
4Qwen 3 Next 80B A3B25.8
5Gemini 2.0 Flash23.1

Interactive version: theaggregate.ai/benchmark?slug=cipher-cross-record-inference-mimic-few-shot · How It Works · Data refreshed daily, snapshot 2026-09-26.