AA Omniscience — leaderboard
AA-Omniscience: factual recall and hallucination rates across 6,000 questions in Law, Health, Business, SWE, Humanities, and Science/Math.
Metric: Score. Source: artificialanalysis.ai. Status: saturated. 449 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3.1 Pro (Preview) | 32.93 |
| 2 | Claude Opus 4.8 (Adaptive Reasoning, Max Effort) | 27.43 |
| 3 | Grok 4.5 (High) | 26.38 |
| 4 | Claude Opus 4.7 (Adaptive Reasoning, Max Effort) | 26.17 |
| 5 | Gemini 3.5 Flash (High) | 22.68 |
| 6 | GPT-5.6 Sol (Max) | 21.7 |
| 7 | Gemini 3.5 Flash (Medium) | 21.65 |
| 8 | GPT-5.6 Sol (xHigh) | 20.55 |
| 9 | GPT-5.5 (xHigh) | 20.07 |
| 10 | GPT-5.6 Sol (High) | 19.75 |
| 11 | GPT-5.6 Sol (Medium) | 18.95 |
| 12 | Kimi K3 | 18.42 |
| 13 | GPT-5.6 Sol (Low) | 18.4 |
| 14 | GPT-5.5 (High) | 18.37 |
| 15 | Grok 4.3 (High) | 18.32 |
Interactive version: theaggregate.ai/benchmark?slug=aa-omniscience · How the rankings work · Data refreshed daily, snapshot 2026-07-22.