SAHARA - Reading Comprehension: leaderboard
Metric: Accuracy (%). Source: huggingface.co. 47 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Pro (Preview) | 58.89 |
| 2 | Gemini 2.5 Pro | 55.92 |
| 3 | GPT-5 | 51.01 |
| 4 | GPT-5.6 Luna | 51 |
| 5 | GPT-5.1 | 50.24 |
| 6 | Gemini 2.5 Flash | 48.01 |
| 7 | Claude Sonnet 4 | 42.33 |
| 8 | GPT-4.1 | 39.7 |
| 9 | Qwen 3.6 35B A3B | 39.44 |
| 10 | Qwen 3.5 27B | 37.57 |
| 11 | Qwen 3.6 27B | 36.96 |
| 12 | Claude Sonnet 4 (20250514) | 36.6 |
| 13 | Llama 3.3 70B Instruct | 35.92 |
| 14 | DeepSeek R1 Distill Llama 70B | 33.72 |
| 15 | DeepSeek R1 Distill Qwen 32B | 33.64 |
Interactive version: theaggregate.ai/benchmark?slug=sahara-reading-comprehension · How It Works · Data refreshed daily, snapshot 2026-09-19.