GPT-3.5 Turbo — benchmark results
OpenAI GPT-3.5 Turbo chat model, tracked as a legacy chat/API baseline. Provider: OpenAI. Released 2022-11-30. Access: API.
Unified ELO 1460 ± 9, rank #955 of 1776 rated models, from 422 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| RABBITS MedMCQA G2B | 97.7 | Accuracy (%) | 100 |
| RABBITS MedMCQA Original | 98.28 | Accuracy (%) | 100 |
| RABBITS MedQA G2B | 96.03 | Accuracy (%) | 98 |
| RABBITS MedQA Original | 96.3 | Accuracy (%) | 96 |
| HumanLikeness - Word-1 | 68.9 | Humanlike Score (%) | 94.7 |
| URIAL-Bench - Coding | 6.9 | Judge Score (0-10) | 94.4 |
| URIAL-Bench - Extraction | 8.85 | Judge Score (0-10) | 94.4 |
| URIAL-Bench - Math | 6.3 | Judge Score (0-10) | 94.4 |
| URIAL-Bench - Overall | 7.94 | Judge Score (0-10) | 94.4 |
| URIAL-Bench - Roleplay | 8.4 | Judge Score (0-10) | 94.4 |
| URIAL-Bench - Turn 1 | 8.07 | Judge Score (0-10) | 94.4 |
| URIAL-Bench - Turn 2 | 7.81 | Judge Score (0-10) | 94.4 |
Interactive version: theaggregate.ai/model?slug=gpt-3-5-turbo · How the rankings work · Data refreshed daily, snapshot 2026-07-22.