HELM Long Context - OpenAI MRCR — leaderboard

Metric: MRCR Accuracy. Source: crfm.stanford.edu. 11 models tracked.

Top models

#ModelScore
1Palmyra X525.62
2Gemini 2.0 Flash21.6
3Llama 4 Maverick Instruct FP821.49
4GPT-4.1 (2025-04-14)21.45
5GPT-4.1 Mini20.79
6Gemini 2.0 Flash Lite17.91
7Llama 4 Scout Instruct17.12
8GPT-4.1 Nano17.04
9Nova Lite11.09
10Nova Pro9.89

Interactive version: theaggregate.ai/benchmark?slug=helm-long-context-openai-mrcr · How the rankings work · Data refreshed daily, snapshot 2026-07-22.