GPT-4.1 (2025-04-14): benchmark results

April 14, 2025 GPT-4.1 snapshot, tracked when sources report the dated API model. Provider: OpenAI. Released 2025-04-14. Access: API.

Unified ELO 1622 ± 1, rank #198 of 1392 rated models, from 486 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HELM Long Context - InfiniteBench En.MC97EM100
HELM Long Context - RULER HotPotQA70RULER String Match100
HELM Long Context - RULER SQuAD88RULER String Match100
MIA-Bench94.55Overall Score100
MedAgentGym70.15Average Score (self-reported)100
Open LMM Reasoning - DynaMath - Subject-statistics; Avg79.6Accuracy (%)100
Open LMM Reasoning - MathVision - analytic geometry58.3Accuracy (%)100
Open LMM Reasoning - MathVision - arithmetic66.4Accuracy (%)100
OpenVLM LLaVA-Bench - Detail103.7Score100
OpenVLM MMBench V1.1 EN - Action Recognition96.6Accuracy (%)100
OpenVLM MMT-Bench - Clock Reading100Score (%)100
OpenVLM MMT-Bench - Font Recognition65Score (%)100

Interactive version: theaggregate.ai/model?slug=gpt-4-1-2025-04-14 · How It Works · Data refreshed daily, snapshot 2026-09-05.