GPT-4o Mini (2024-07-18): benchmark results

July 18, 2024 GPT-4o Mini snapshot, tracked when sources report the dated API model. Provider: OpenAI. Released 2024-07-18. Access: API.

Unified ELO 1517 ± 1, rank #615 of 1392 rated models, from 356 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HELM MedHELM - EHRSHOT90.67EM100
KOFFVQA - Korean OCR100Score (%)100
LLM Trustworthy - Adversarial Demo88.49Trust Score (%)100
Thai LLM - Writing8.35Rating (0-10)100
Thai LLM - Knowledge III6.45Rating (0-10)98.6
PIQA88.7Accuracy (%)98.3
Thai LLM - STEM7.75Rating (0-10)97.2
Thai LLM NLU - xnli.tha_seacrowd_pairs52.06Accuracy (%)97.1
Thai LLM NLG - flores200_eng_Latn_tha_Thai_seacrowd_t2t (BLEU)30.44BLEU97
Thai LLM NLG - flores200_eng_Latn_tha_Thai_seacrowd_t2t (SacreBLEU)42.02SacreBLEU97
Thai LLM NLG - flores200_eng_Latn_tha_Thai_seacrowd_t2t (chrF++)56.92chrF++97
WildBench57.14WB Score Task-Macro96.8

Interactive version: theaggregate.ai/model?slug=gpt-4o-mini-2024-07-18 · How It Works · Data refreshed daily, snapshot 2026-09-05.