Grok 3 Mini (High) — benchmark results
Grok 3 Mini evaluated at the high reasoning-effort setting. Provider: xAI. Released 2025-02-17. Access: API.
Unified ELO 1692 ± 18, rank #287 of 1776 rated models, from 33 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| MATH-Perturb (Hard) | 87.1 | Accuracy (%) | 94.4 |
| BenchTable | 67.3 | Total Score (%) | 82.8 |
| Confabulation Leaderboard (Lechmazur) | 6.93 | Confabulation rate % (lower is better) | 82.5 |
| Vals AI MATH 500 | 94.2 | Accuracy (%) | 80.9 |
| LLM Chess (Saplin) | 430.4 | ELO | 80.7 |
| Step Game (Lechmazur) | 3.24 | TrueSkill μ | 78.4 |
| MATH Level 5 | 88.07 | Accuracy (%) | 75.9 |
| PlatinumBench (MIT) | 1.51 | Avg Error Rate (%) | 69.7 |
| Vals AI LegalBench | 83.14 | Accuracy (%) | 68 |
| Vals AI AIME | 85 | Accuracy (%) | 61.1 |
| OTIS Mock AIME 2024-25 | 77.78 | Accuracy (%) | 60.9 |
| LiveCodeBench | 78.1 | Pass@1 avg (%) | 59.3 |
Interactive version: theaggregate.ai/model?slug=grok-3-mini-high · How the rankings work · Data refreshed daily, snapshot 2026-07-22.