Mercury 2: benchmark results
Inception's Mercury 2 diffusion LLM, generating via parallel refinement at over 1,000 tokens/s. Provider: Inception. Released 2026-02-24. Access: API.
Unified ELO 1595 ± 1, rank #276 of 1392 rated models, from 60 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| PinchBench | 100 | Success Rate (%) | 99.1 |
| CritPt | 80 | Accuracy (self-reported) | 95.1 |
| Story Theory Bench | 99.1 | Score (%) | 94.1 |
| AA IFBench | 69.8 | Accuracy (%) | 84 |
| SnakeBench | 28.1 | TrueSkill Rating | 80.8 |
| AA Humanity's Last Exam | 17.15 | Accuracy (%) | 68.6 |
| AA Terminal-Bench Hard | 26.52 | Accuracy (%) | 68.1 |
| AA TAU-2 Bench | 70.76 | Accuracy (%) | 63.3 |
| AA Omniscience - Health | 24 | Accuracy (%) | 62.8 |
| AA GPQA Diamond | 76.97 | Accuracy (%) | 61.4 |
| AA Omniscience - Science, Engineering & Mathematics | 31.19 | Accuracy (%) | 61.3 |
| AA CritPt | 0.85 | Accuracy (%) | 60.7 |
Interactive version: theaggregate.ai/model?slug=mercury-2 · How It Works · Data refreshed daily, snapshot 2026-09-05.