Grok 3 Mini (Low) — benchmark results
Grok 3 Mini evaluated at the low reasoning-effort setting. Provider: xAI. Released 2025-02-17. Access: API.
Unified ELO 1582 ± 31, rank #527 of 1776 rated models, from 25 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| MATH Level 5 | 90.94 | Accuracy (%) | 80.6 |
| Confabulation Leaderboard (Lechmazur) | 10.89 | Confabulation rate % (lower is better) | 77 |
| Vals AI TaxEval v2 | 72.98 | Accuracy (%) | 64.8 |
| Generalization V1 (Lechmazur) | 1.9 | Avg Rank (lower is better) | 55.6 |
| LLM Chess (Saplin) | -33.9 | ELO | 55 |
| Vals AI LegalBench | 81.92 | Accuracy (%) | 54.4 |
| NYT Connections Older Models | 26 | Score (%) | 51.4 |
| OTIS Mock AIME 2024-25 | 62.22 | Accuracy (%) | 49.4 |
| Vals AI CorpFin v2 | 59.48 | Accuracy (%) | 44.3 |
| Vals AI AIME | 70.62 | Accuracy (%) | 44.2 |
| Vals AI MGSM | 90.36 | Accuracy (%) | 44 |
| Vals AI MedQA | 88.65 | Accuracy (%) | 43.6 |
Interactive version: theaggregate.ai/model?slug=grok-3-mini-low · How the rankings work · Data refreshed daily, snapshot 2026-07-22.