O4 Mini Deep Research: benchmark results
Provider: OpenAI. Access: API.
Unified ELO 1762 ± 26, rank #100 of 1629 rated models, from 17 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Mr Dre - Content Feedback Incorporation Rate | 93.6 | Incorporation rate (%; turn-2 revision after content feedbac | 100 |
| Mr Dre - Format Feedback Incorporation Rate | 98.5 | Incorporation rate (%; turn-2 revision after format feedback | 100 |
| DailyReport (Deep Research Agents) - Factuality | 66.3 | Factuality (0-1 scaled to %), the share of extracted objecti | 75 |
| DailyReport (Deep Research Agents) - SubTask Pass | 24.1 | Share of subtasks (0-1 scaled to %) that satisfy every rubri | 75 |
| Mr Dre - Content Feedback Break Rate | 29.7 | Break rate (%; share of checklist criteria covered in the tu | 75 |
| Mr Dre - Format Feedback Break Rate | 14.8 | Break rate (%; share of checklist criteria covered in the tu | 75 |
| MedProbeBench (Deep Research Agents) - Claim Recall | 44.4 | Task success rate (x100): share of the expert guideline atom | 66.7 |
| SnakeBench | 23.6 | TrueSkill Rating | 57.2 |
| DailyReport (Deep Research Agents) - Rationality | 77.8 | Rationality (0-1 scaled to %), rubric scores of 0, 0.5 or 1 | 50 |
| MedProbeBench (Deep Research Agents) | 56.1 | Overall score (x100): task-weighted composite of holistic qu | 50 |
| MedProbeBench (Deep Research Agents) - Evidence Verification | 50.2 | Evidence verification (x100): task-weighted composite of cla | 50 |
| MedProbeBench (Deep Research Agents) - Holistic Quality | 69.9 | Holistic quality (x100): task-weighted mean of comprehensive | 50 |
Interactive version: theaggregate.ai/model?slug=o4-mini-deep-research · How It Works · Data refreshed daily, snapshot 2026-10-07.