GPT-4.1: benchmark results
OpenAI's GPT-4.1 model with improved coding, instruction following, and a 1M-token context. Provider: OpenAI. Released 2025-04-14. Access: API.
Unified ELO 1618 ± 1, rank #213 of 1392 rated models, from 1105 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AGC-Bench - futuregen | 2.15 | Dataset z-score | 100 |
| ArabicCulturalQA | 79.1 | MCQ Accuracy Avg (%) | 100 |
| AsyncTool | 38.06 | Overall (self-reported) | 100 |
| BlueBench | 64.61 | Average Score (%) | 100 |
| EuroEval Hungarian Summarization - Hunsum | 34.8 | Score (%) | 100 |
| EuroEval Romanian Summarization - Sumo RO | 38.27 | Score (%) | 100 |
| Galileo Agent - Banking Accuracy | 60 | Accuracy (%) | 100 |
| Galileo Agent - Investment Accuracy | 64 | Accuracy (%) | 100 |
| Galileo Agent Leaderboard | 62 | Avg Accuracy (%) | 100 |
| KOFFVQA - Object Attributes | 88.33 | Score (%) | 100 |
| Large Language Models Lack Temporal Awareness | 71.11 | Accuracy (self-reported) | 100 |
| MERA Code - RealCodeJava | 41.61 | pass@1 (%) | 100 |
Interactive version: theaggregate.ai/model?slug=gpt-4-1 · How It Works · Data refreshed daily, snapshot 2026-09-05.