GPT-4.1 (2025-04-14) — benchmark results
April 14, 2025 GPT-4.1 snapshot, tracked when sources report the dated API model. Provider: OpenAI. Released 2025-04-14. Access: API.
Unified ELO 1625 ± 8, rank #422 of 1776 rated models, from 496 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HELM Long Context - InfiniteBench En.MC | 97 | EM | 100 |
| HELM Long Context - RULER HotPotQA | 70 | RULER String Match | 100 |
| HELM Long Context - RULER SQuAD | 88 | RULER String Match | 100 |
| MIA-Bench | 94.55 | Overall Score | 100 |
| Open LMM Reasoning - DynaMath - Subject-statistics; Avg | 79.6 | Accuracy (%) | 100 |
| Open LMM Reasoning - MathVision - analytic geometry | 58.3 | Accuracy (%) | 100 |
| Open LMM Reasoning - MathVision - arithmetic | 66.4 | Accuracy (%) | 100 |
| Open LMM Reasoning - MathVista - LOG | 54.1 | Accuracy (%) | 100 |
| OpenVLM LLaVA-Bench - Detail | 103.7 | Score | 100 |
| OpenVLM MMBench V1.1 EN - Action Recognition | 96.6 | Accuracy (%) | 100 |
| OpenVLM MMT-Bench - Clock Reading | 100 | Score (%) | 100 |
| OpenVLM MMT-Bench - Font Recognition | 65 | Score (%) | 100 |
Interactive version: theaggregate.ai/model?slug=gpt-4-1-2025-04-14 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.