Grok Beta: benchmark results

xAI's grok-beta, the Grok-2-class 128K-context preview model that launched the public xAI API beta in November 2024. Provider: xAI. Released 2024-11-04. Access: API.

Unified ELO 1534 ± 1, rank #519 of 1392 rated models, from 29 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
SpeechMap Compliance93.7% Requests Completed96.2
TextClass Benchmark1741.94Meta-Elo (self-reported)91.6
ProLLM - Function Calling91.8Score (%)90.7
VNTL Leaderboard71.27Accuracy (%)89.5
EvalPlus (HumanEval+ & MBPP+)73Pass@1 avg (%)86.3
ProLLM - Summarization74.1Score (%)61.2
ForecastBench64.3Overall Score (higher is better)61
AI for Education Pedagogy - Technology80.19Accuracy (%)57.4
HumanEval+80.5HumanEval+ pass@1 (self-reported)56.5
ProLLM - StackEval90.4Score (%)51.2
NYT Connections Original23.7Score (%)48.3
ProLLM - Q&A Assistant93.1Score (%)44.1

Interactive version: theaggregate.ai/model?slug=grok-beta · How It Works · Data refreshed daily, snapshot 2026-09-05.