GPT-4o Mini (2024-07-18): benchmark results
July 18, 2024 GPT-4o Mini snapshot, tracked when sources report the dated API model. Provider: OpenAI. Released 2024-07-18. Access: API.
Unified ELO 1517 ± 1, rank #615 of 1392 rated models, from 356 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HELM MedHELM - EHRSHOT | 90.67 | EM | 100 |
| KOFFVQA - Korean OCR | 100 | Score (%) | 100 |
| LLM Trustworthy - Adversarial Demo | 88.49 | Trust Score (%) | 100 |
| Thai LLM - Writing | 8.35 | Rating (0-10) | 100 |
| Thai LLM - Knowledge III | 6.45 | Rating (0-10) | 98.6 |
| PIQA | 88.7 | Accuracy (%) | 98.3 |
| Thai LLM - STEM | 7.75 | Rating (0-10) | 97.2 |
| Thai LLM NLU - xnli.tha_seacrowd_pairs | 52.06 | Accuracy (%) | 97.1 |
| Thai LLM NLG - flores200_eng_Latn_tha_Thai_seacrowd_t2t (BLEU) | 30.44 | BLEU | 97 |
| Thai LLM NLG - flores200_eng_Latn_tha_Thai_seacrowd_t2t (SacreBLEU) | 42.02 | SacreBLEU | 97 |
| Thai LLM NLG - flores200_eng_Latn_tha_Thai_seacrowd_t2t (chrF++) | 56.92 | chrF++ | 97 |
| WildBench | 57.14 | WB Score Task-Macro | 96.8 |
Interactive version: theaggregate.ai/model?slug=gpt-4o-mini-2024-07-18 · How It Works · Data refreshed daily, snapshot 2026-09-05.