Grok 3 — benchmark results

xAI's Grok 3 flagship model for reasoning and general tasks (February 2025). Provider: xAI. Released 2025-02-17. Access: API.

Unified ELO 1601 ± 10, rank #475 of 1776 rated models, from 269 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
UGI Leaderboard63.47UGI Score99.3
ProLLM - Entity Extraction85.5Score (%)97.9
SpeechMap Compliance96.1% Requests Completed97.6
EuroEval Swedish NLU - ScaLA SV71.91Linguistic acceptability Score (%)97.3
EuroEval Icelandic NLU - Hotter and Colder Sentiment58.9Sentiment classification Score (%)97.2
EuroEval Norwegian Knowledge71.26Knowledge Average Score (%)96.5
EuroEval Finnish Common Sense Reasoning80.26Common Sense Reasoning Average Score (%)96.4
EuroEval Portuguese Common Sense Reasoning93.05Common Sense Reasoning Average Score (%)96.4
EuroEval Danish73.34Average Score (%)96.1
EuroEval Icelandic56.5Average Score (%)95.4
EuroEval French NLU - Allocine96.47Sentiment classification Score (%)95.3
EuroEval Danish Knowledge93.1Knowledge Average Score (%)95.1

Interactive version: theaggregate.ai/model?slug=grok-3 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.