GPT-3.5 Turbo Instruct: benchmark results
Provider: OpenAI. Released 2022-11-30. Access: API.
Unified ELO 1481 ± 48, rank #1273 of 2656 rated models, from 13 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AI Chess Leaderboard (Continuation) | 1436 | Elo | 94.4 |
| Chess Bench LLM | 1586 | Lichess Rating | 91.2 |
| LM Market Cap LMC Score | 40 | LMC Score (0-100) | 28 |
| Klu LLM Leaderboard | 70 | Klu Index | 27.5 |
| AI Chess Leaderboard (Reasoning) | 553 | Elo | 22.8 |
| AgentDrive - Physics | 27.5 | Accuracy (%) | 22 |
| AgentDrive - Scenario | 82.5 | Accuracy (%) | 12 |
| METR Benchmark | 0.01 | 50% Time Horizon (hours) | 8 |
| AgentDrive - Comparative | 52.5 | Accuracy (%) | 7 |
| AgentDrive - Hybrid | 5 | Accuracy (%) | 7 |
| AgentDrive | 34.5 | Accuracy (%) | 6 |
| METR Benchmark (80% Horizon) | 0 | 80% Time Horizon (hours) | 4 |
Interactive version: theaggregate.ai/model?slug=gpt-3-5-turbo-instruct · How It Works · Data refreshed daily, snapshot 2026-09-19.