GPT-4.1 — benchmark results
OpenAI's GPT-4.1 model with improved coding, instruction following, and a 1M-token context. Provider: OpenAI. Released 2025-04-14. Access: API.
Unified ELO 1634 ± 5, rank #404 of 1776 rated models, from 1000 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AGC-Bench - futuregen | 2.15 | Dataset z-score | 100 |
| AsyncTool | 38.06 | Overall (self-reported) | 100 |
| BlueBench | 64.61 | Average Score (%) | 100 |
| EuroEval Hungarian Summarization - Hunsum | 34.8 | Score (%) | 100 |
| EuroEval Romanian Summarization - Sumo RO | 38.27 | Score (%) | 100 |
| Galileo Agent - Banking Accuracy | 60 | Accuracy (%) | 100 |
| Galileo Agent - Investment Accuracy | 64 | Accuracy (%) | 100 |
| Galileo Agent Leaderboard | 62 | Avg Accuracy (%) | 100 |
| KOFFVQA - Object Attributes | 88.33 | Score (%) | 100 |
| Large Language Models Lack Temporal Awareness | 71.11 | Accuracy (self-reported) | 100 |
| MHGraphBench | 70.28 | Avg$_{All^\ast}$ (self-reported) | 100 |
| MedAgentGym | 70.15 | Average Score (self-reported) | 100 |
Interactive version: theaggregate.ai/model?slug=gpt-4-1 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.