GPT-4o Search Preview: benchmark results
Provider: OpenAI. Released 2025-03-11. Access: API.
Unified ELO 1690 ± 27, rank #350 of 2656 rated models, from 22 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Tinybird AI SQL Benchmark - Exactness | 57.21 | Result exactness vs human reference queries (0-100) | 94.8 |
| DeepResearch Bench - Citation Accuracy | 86.63 | Score (%) | 91.7 |
| LLMEval-Fair - Medicine | 8.27 | Discipline Score (0-10) | 78.4 |
| LLMEval-Fair - Literature | 7.77 | Discipline Score (0-10) | 73.3 |
| Tinybird AI SQL Benchmark - Success Rate | 100 | Questions answered with a valid query within 3 retries (%) | 72.5 |
| LLMEval-Fair - Education | 8.43 | Discipline Score (0-10) | 67.2 |
| LLMEval-Fair - Engineering | 8.27 | Discipline Score (0-10) | 67.2 |
| LLMEval-Fair - History | 8.73 | Discipline Score (0-10) | 67.2 |
| LLMEval-Fair - Law | 8.67 | Discipline Score (0-10) | 67.2 |
| LLMEval-Fair - Overall | 83.73 | Absolute Score (0-100) | 67.2 |
| LLMEval-Fair - Economics | 8.77 | Discipline Score (0-10) | 65.5 |
| LLMEval-Fair - Management | 8.8 | Discipline Score (0-10) | 63.8 |
Interactive version: theaggregate.ai/model?slug=gpt-4o-search-preview · How It Works · Data refreshed daily, snapshot 2026-09-19.