Grok 4.1 (Thinking) — benchmark results
Provider: xAI. Released 2025-11-17. Access: API.
Unified ELO 1695 ± 79, rank #337 of 1806 rated models, from 7 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Chatbot Arena (Text) | 1466 | Elo | 89.8 |
| CyBench | 39 | End-to-End % Solved | 62.5 |
| LLM Stats Score | 16.52 | LLM Stats Score (conservative rating) | 42.1 |
| PrinzBench | 25 | Score (x/99) | 28.6 |
| Creative Writing v3 | 86.09 | Elo score (self-reported) | 7.4 |
| Chatbot Arena (Code) | 1210 | Elo | 5.5 |
| FigQA | 34 | Score (self-reported) | 0 |
Interactive version: theaggregate.ai/model?slug=grok-4-1-thinking · How It Works · Data refreshed daily, snapshot 2026-08-07.