AA-Omniscience Net Score: leaderboard

AA-Omniscience factuality results reported by Anthropic as net score: correct responses minus incorrect responses, with abstentions scoring zero.

Metric: Net score (self-reported). Source: benchmarklist.com. Status: years away from saturation. 10 models tracked.

Top models

#ModelScore
1Claude Mythos 558
2Claude Mythos Preview54
3Claude Opus 549
4Claude Opus 4.841
5Claude Opus 4.738
6Claude Sonnet 523
7Claude Opus 4.621
8Claude Opus 4.515
9Claude Sonnet 4.614

Interactive version: theaggregate.ai/benchmark?slug=aa-omniscience-net-score · How It Works · Data refreshed daily, snapshot 2026-09-05.