Grok 4.1 — benchmark results

xAI's Grok 4.1 model with improved conversational and emotional intelligence. Provider: xAI. Released 2025-11-17. Access: API.

Unified ELO 1693 ± 39, rank #285 of 1776 rated models, from 20 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
SnitchBench92.5Govt Contact Rate (%)100
LiveMedBench28.28Overall Score (%)91.9
Chatbot Arena (Text)1466Elo90.8
OpenCompass Agent - Multi-Turn80.5Score (%)84
BenchLM60Overall Score75.6
OpenCompass Language - Dialogue96.2Score (%)56
OpenCompass Code - Competition42.7Score (%)44
Ducky Bench (Stabby Quack)892ELO39.3
PM-LLM-Benchmark28.9Score33
Mercor APEX12.8Best Pass@1 (%)29.2
PrinzBench25Score (x/99)28.6
TrackingAI IQ Test (Mensa Norway)51.43Mensa Norway Score (%)27.3

Interactive version: theaggregate.ai/model?slug=grok-4-1 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.