MIRA-Math: leaderboard
Metric: Final-answer accuracy (%; exact verifier match; zero-shot; 2,310 generated instances from 22 typed families, one hidden atomic fact per instance, request budget 3/6/9 by difficulty, fixed gpt-4o-mini responder, temperature 0, single pass; fraction x100). Source: arxiv.org. Saturation forecast: Around December 2026. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemma 4 31B (IT) | 73.9 |
| 2 | GPT-5.1 | 69 |
| 3 | Grok 4.20 | 50.3 |
| 4 | Gemini 2.5 Flash | 49 |
| 5 | GPT-4o Mini | 24.8 |
Interactive version: theaggregate.ai/benchmark?slug=mira-math · How It Works · Data refreshed daily, snapshot 2026-09-29.