HELM Long Context - InfiniteBench En.Sum — leaderboard

Metric: ROUGE-L. Source: crfm.stanford.edu. 11 models tracked.

Top models

#ModelScore
1Llama 4 Scout Instruct17.6
2GPT-4.1 (2025-04-14)17.38
3Llama 4 Maverick Instruct FP816.06
4GPT-4.1 Mini16.05
5Gemini 2.0 Flash Lite15.51
6Gemini 2.0 Flash15.11
7Nova Lite14.76
8Palmyra X514.65
9GPT-4.1 Nano11.34
10Nova Pro10.95

Interactive version: theaggregate.ai/benchmark?slug=helm-long-context-infinitebench-en-sum · How the rankings work · Data refreshed daily, snapshot 2026-07-22.