Grok 4.1 Fast: benchmark results
xAI's speed-oriented Grok 4.1 variant for agentic tool-calling workloads with a 2M-token context. Provider: xAI. Released 2025-11-19. Access: API.
Unified ELO 1626 ± 1, rank #190 of 1392 rated models, from 174 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Hack-Verifiable TextArena | 28.5 | Avg HR (self-reported) | 100 |
| Ko-AgentBench - L5 Error Handling & Robustness | 34.75 | Adaptive Routing Score (%) | 100 |
| ALL Bench LLM | 81.77 | Average Numeric Benchmark Score (%) | 97.4 |
| MCP-Universe (LLM w/ Function Calls) | 60.58 | Avg Evaluator Score | 95 |
| RAI-Bench - RAG Robustness (LC Abstention) | 95 | Rate (%) | 94.9 |
| FormationEval | 97.6 | Accuracy (%) | 94.4 |
| MLX Benchmark V2 - Coding | 63.64 | Accuracy (%) | 90 |
| Pencil Puzzle Bench - Hitori | 26.7 | Direct-ask Success Rate (%) | 89 |
| Pencil Puzzle Bench - Nurimaze | 6.7 | Direct-ask Success Rate (%) | 89 |
| GSMA Open-Telco - TeleLogs | 71 | Score (%) | 88.6 |
| Chess Bench LLM | 1400 | Lichess Rating | 87.7 |
| RAI-Bench - RAG Robustness (LC Factuality) | 47 | Rate (%) | 87 |
Interactive version: theaggregate.ai/model?slug=grok-4-1-fast · How It Works · Data refreshed daily, snapshot 2026-09-05.