Grok 3 Mini — benchmark results
xAI's lightweight Grok 3 reasoning model that thinks before responding, aimed at cost-efficient math, logic and coding (February 2025). Provider: xAI. Released 2025-02-17. Access: API.
Unified ELO 1588 ± 13, rank #507 of 1776 rated models, from 170 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| MMTU - Table Join | 77.14 | Accuracy (%) | 100 |
| EuroEval Norwegian Common Sense Reasoning | 90.07 | Common Sense Reasoning Average Score (%) | 98.6 |
| LLM Stats (AIME 2024) | 95.8 | Score (%) | 98.1 |
| MMTU - Column Relationship | 67.88 | Accuracy (%) | 96 |
| ProLLM - OpenBook Q&A | 86.7 | Score (%) | 95.7 |
| EuroEval Faroese NLU - ScaLA FO | 34.07 | Linguistic acceptability Score (%) | 95.6 |
| EuroEval Icelandic Common Sense Reasoning | 76.18 | Common Sense Reasoning Average Score (%) | 95.4 |
| EuroEval Italian Knowledge | 83.74 | Knowledge Average Score (%) | 95.4 |
| ProLLM - Function Calling | 92.3 | Score (%) | 95.4 |
| EuroEval Portuguese Knowledge | 87.04 | Knowledge Average Score (%) | 95.2 |
| EuroEval Spanish Knowledge | 82.35 | Knowledge Average Score (%) | 95 |
| EuroEval German NLU - Sb10K | 60.13 | Sentiment classification Score (%) | 94.5 |
Interactive version: theaggregate.ai/model?slug=grok-3-mini · How the rankings work · Data refreshed daily, snapshot 2026-07-22.