InfiniteBench — leaderboard
InfiniteBench: Measures long-context retrieval, needle finding, summarization, factual grounding, or retrieval-augmented generation quality.
Metric: Avg Score (%). Source: github.com. Status: years away from saturation. 7 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-4 | 45.89 |
| 2 | Claude 2 | 37.22 |
| 3 | Yi-34B-200K | 26.36 |
| 4 | Yi-6B-200K | 23.33 |
Interactive version: theaggregate.ai/benchmark?slug=infinitebench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.