GPT-3.5 Turbo 16K: benchmark results

Provider: OpenAI. Released 2022-11-30. Access: API.

Unified ELO 1460 ± 24, rank #1441 of 2656 rated models, from 21 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
CLAMBER - Few-shot - Clarification BERTScore31.16BERTScore (0-100)100
CLAMBER - Few-shot - Clarification Helpfulness51.58Helpfulness (%)100
CLAMBER - Few-shot - Identification Accuracy51.66Accuracy (%)100
CLAMBER - Few-shot - Identification F149.28F1 (%)100
CLAMBER - Few-shot CoT - Clarification BERTScore33.48BERTScore (0-100)100
CLAMBER - Few-shot CoT - Clarification Helpfulness53.29Helpfulness (%)100
CLAMBER - Zero-shot - Clarification BERTScore27.47BERTScore (0-100)100
CLAMBER - Zero-shot - Clarification Helpfulness46.45Helpfulness (%)100
CLAMBER - Zero-shot - Identification F153.45F1 (%)100
CLAMBER - Zero-shot CoT - Clarification BERTScore30.22BERTScore (0-100)100
CLAMBER - Zero-shot CoT - Clarification Helpfulness50.47Helpfulness (%)100
CLAMBER - Zero-shot CoT - Identification Accuracy57.38Accuracy (%)100

Interactive version: theaggregate.ai/model?slug=gpt-3-5-turbo-16k · How It Works · Data refreshed daily, snapshot 2026-09-19.