Grok 3 Mini — benchmark results

xAI's lightweight Grok 3 reasoning model that thinks before responding, aimed at cost-efficient math, logic and coding (February 2025). Provider: xAI. Released 2025-02-17. Access: API.

Unified ELO 1588 ± 13, rank #507 of 1776 rated models, from 170 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
MMTU - Table Join77.14Accuracy (%)100
EuroEval Norwegian Common Sense Reasoning90.07Common Sense Reasoning Average Score (%)98.6
LLM Stats (AIME 2024)95.8Score (%)98.1
MMTU - Column Relationship67.88Accuracy (%)96
ProLLM - OpenBook Q&A86.7Score (%)95.7
EuroEval Faroese NLU - ScaLA FO34.07Linguistic acceptability Score (%)95.6
EuroEval Icelandic Common Sense Reasoning76.18Common Sense Reasoning Average Score (%)95.4
EuroEval Italian Knowledge83.74Knowledge Average Score (%)95.4
ProLLM - Function Calling92.3Score (%)95.4
EuroEval Portuguese Knowledge87.04Knowledge Average Score (%)95.2
EuroEval Spanish Knowledge82.35Knowledge Average Score (%)95
EuroEval German NLU - Sb10K60.13Sentiment classification Score (%)94.5

Interactive version: theaggregate.ai/model?slug=grok-3-mini · How the rankings work · Data refreshed daily, snapshot 2026-07-22.