GPT-3.5 Turbo 16K: benchmark results
Provider: OpenAI. Released 2022-11-30. Access: API.
Unified ELO 1460 ± 24, rank #1441 of 2656 rated models, from 21 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| CLAMBER - Few-shot - Clarification BERTScore | 31.16 | BERTScore (0-100) | 100 |
| CLAMBER - Few-shot - Clarification Helpfulness | 51.58 | Helpfulness (%) | 100 |
| CLAMBER - Few-shot - Identification Accuracy | 51.66 | Accuracy (%) | 100 |
| CLAMBER - Few-shot - Identification F1 | 49.28 | F1 (%) | 100 |
| CLAMBER - Few-shot CoT - Clarification BERTScore | 33.48 | BERTScore (0-100) | 100 |
| CLAMBER - Few-shot CoT - Clarification Helpfulness | 53.29 | Helpfulness (%) | 100 |
| CLAMBER - Zero-shot - Clarification BERTScore | 27.47 | BERTScore (0-100) | 100 |
| CLAMBER - Zero-shot - Clarification Helpfulness | 46.45 | Helpfulness (%) | 100 |
| CLAMBER - Zero-shot - Identification F1 | 53.45 | F1 (%) | 100 |
| CLAMBER - Zero-shot CoT - Clarification BERTScore | 30.22 | BERTScore (0-100) | 100 |
| CLAMBER - Zero-shot CoT - Clarification Helpfulness | 50.47 | Helpfulness (%) | 100 |
| CLAMBER - Zero-shot CoT - Identification Accuracy | 57.38 | Accuracy (%) | 100 |
Interactive version: theaggregate.ai/model?slug=gpt-3-5-turbo-16k · How It Works · Data refreshed daily, snapshot 2026-09-19.