GPT-4.1: benchmark results

OpenAI's GPT-4.1 model with improved coding, instruction following, and a 1M-token context. Provider: OpenAI. Released 2025-04-14. Access: API.

Unified ELO 1618 ± 1, rank #213 of 1392 rated models, from 1105 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AGC-Bench - futuregen2.15Dataset z-score100
ArabicCulturalQA79.1MCQ Accuracy Avg (%)100
AsyncTool38.06Overall (self-reported)100
BlueBench64.61Average Score (%)100
EuroEval Hungarian Summarization - Hunsum34.8Score (%)100
EuroEval Romanian Summarization - Sumo RO38.27Score (%)100
Galileo Agent - Banking Accuracy60Accuracy (%)100
Galileo Agent - Investment Accuracy64Accuracy (%)100
Galileo Agent Leaderboard62Avg Accuracy (%)100
KOFFVQA - Object Attributes88.33Score (%)100
Large Language Models Lack Temporal Awareness71.11Accuracy (self-reported)100
MERA Code - RealCodeJava41.61pass@1 (%)100

Interactive version: theaggregate.ai/model?slug=gpt-4-1 · How It Works · Data refreshed daily, snapshot 2026-09-05.