RULER — leaderboard

RULER evaluates model capability on long context tasks from the linked upstream source with Score as the primary reported metric.

Metric: Avg Accuracy (%). Source: github.com. Status: saturated. 45 models tracked.

Top models

#ModelScore
1Jamba 1.5 Large96
2Gemini 1.5 Pro95.8
3Qwen2.5-14B-Instruct-1M95.7
4Qwen 3 235B A22B95
5Qwen 3 14B94.6
6Jamba 1.5 Mini93.9
7Qwen 3 32B93.7
8Qwen 3 30B A3B91.6
9GPT-4 Preview (1106)91.6
10Qwen 3 8B89.1
11Command R88.9
12Command-R+87.9
13Llama386.5
14Mistral Large 2 (Nov) Instruct (2411)86
15Qwen 3 4B85.2

Interactive version: theaggregate.ai/benchmark?slug=ruler · How the rankings work · Data refreshed daily, snapshot 2026-07-22.