Grok 4.1 — benchmark results
xAI's Grok 4.1 model with improved conversational and emotional intelligence. Provider: xAI. Released 2025-11-17. Access: API.
Unified ELO 1693 ± 39, rank #285 of 1776 rated models, from 20 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| SnitchBench | 92.5 | Govt Contact Rate (%) | 100 |
| LiveMedBench | 28.28 | Overall Score (%) | 91.9 |
| Chatbot Arena (Text) | 1466 | Elo | 90.8 |
| OpenCompass Agent - Multi-Turn | 80.5 | Score (%) | 84 |
| BenchLM | 60 | Overall Score | 75.6 |
| OpenCompass Language - Dialogue | 96.2 | Score (%) | 56 |
| OpenCompass Code - Competition | 42.7 | Score (%) | 44 |
| Ducky Bench (Stabby Quack) | 892 | ELO | 39.3 |
| PM-LLM-Benchmark | 28.9 | Score | 33 |
| Mercor APEX | 12.8 | Best Pass@1 (%) | 29.2 |
| PrinzBench | 25 | Score (x/99) | 28.6 |
| TrackingAI IQ Test (Mensa Norway) | 51.43 | Mensa Norway Score (%) | 27.3 |
Interactive version: theaggregate.ai/model?slug=grok-4-1 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.