GPT-4.1 Mini — benchmark results
OpenAI's mid-size GPT-4.1 tier with a 1M-token context, cheaper and lower-latency than GPT-4o at similar quality (April 2025). Provider: OpenAI. Released 2025-04-14. Access: API.
Unified ELO 1547 ± 6, rank #626 of 1776 rated models, from 1062 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| BlueBench - Chatbot Abilities | 97.55 | Score (%) | 100 |
| BlueBench - RAG General | 54.58 | Score (%) | 100 |
| EuroEval Bosnian | 63.93 | Average Score (%) | 100 |
| EuroEval Bosnian NLU - MMS BS | 56.43 | Sentiment classification Score (%) | 100 |
| Galileo Agent - Telecom Accuracy | 64 | Accuracy (%) | 100 |
| Open LMM Reasoning - DynaMath - Level-elementary school | 65.1 | Accuracy (%) | 100 |
| Open LMM Reasoning - DynaMath - Level-elementary school; Avg | 78.1 | Accuracy (%) | 100 |
| Open LMM Reasoning - DynaMath - Level-high school | 47.3 | Accuracy (%) | 100 |
| Open LMM Reasoning - DynaMath - Subject-algebra | 82.4 | Accuracy (%) | 100 |
| Open LMM Reasoning - DynaMath - Subject-algebra; Avg | 95.5 | Accuracy (%) | 100 |
| Open LMM Reasoning - DynaMath - Subject-arithmetic | 57.7 | Accuracy (%) | 100 |
| Open LMM Reasoning - DynaMath - Subject-arithmetic; Avg | 73.8 | Accuracy (%) | 100 |
Interactive version: theaggregate.ai/model?slug=gpt-4-1-mini · How the rankings work · Data refreshed daily, snapshot 2026-07-22.