Mercury 2: benchmark results

Inception's Mercury 2 diffusion LLM, generating via parallel refinement at over 1,000 tokens/s. Provider: Inception. Released 2026-02-24. Access: API.

Unified ELO 1595 ± 1, rank #276 of 1392 rated models, from 60 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
PinchBench100Success Rate (%)99.1
CritPt80Accuracy (self-reported)95.1
Story Theory Bench99.1Score (%)94.1
AA IFBench69.8Accuracy (%)84
SnakeBench28.1TrueSkill Rating80.8
AA Humanity's Last Exam17.15Accuracy (%)68.6
AA Terminal-Bench Hard26.52Accuracy (%)68.1
AA TAU-2 Bench70.76Accuracy (%)63.3
AA Omniscience - Health24Accuracy (%)62.8
AA GPQA Diamond76.97Accuracy (%)61.4
AA Omniscience - Science, Engineering & Mathematics31.19Accuracy (%)61.3
AA CritPt0.85Accuracy (%)60.7

Interactive version: theaggregate.ai/model?slug=mercury-2 · How It Works · Data refreshed daily, snapshot 2026-09-05.