Grok 3 Mini (High) — benchmark results

Grok 3 Mini evaluated at the high reasoning-effort setting. Provider: xAI. Released 2025-02-17. Access: API.

Unified ELO 1692 ± 18, rank #287 of 1776 rated models, from 33 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
MATH-Perturb (Hard)87.1Accuracy (%)94.4
BenchTable67.3Total Score (%)82.8
Confabulation Leaderboard (Lechmazur)6.93Confabulation rate % (lower is better)82.5
Vals AI MATH 50094.2Accuracy (%)80.9
LLM Chess (Saplin)430.4ELO80.7
Step Game (Lechmazur)3.24TrueSkill μ78.4
MATH Level 588.07Accuracy (%)75.9
PlatinumBench (MIT)1.51Avg Error Rate (%)69.7
Vals AI LegalBench83.14Accuracy (%)68
Vals AI AIME85Accuracy (%)61.1
OTIS Mock AIME 2024-2577.78Accuracy (%)60.9
LiveCodeBench78.1Pass@1 avg (%)59.3

Interactive version: theaggregate.ai/model?slug=grok-3-mini-high · How the rankings work · Data refreshed daily, snapshot 2026-07-22.