GPT-NeoXT-Chat-Base-20B — benchmark results
Provider: Together. Released 2023-03-03. Access: Open.
Unified ELO 1168 ± 12, rank #1721 of 1776 rated models, from 78 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| MMLU-by-task - College Chemistry | 39 | Accuracy (%) | 70.2 |
| MMLU-by-task - College Mathematics | 34 | Accuracy (%) | 69.2 |
| MMLU-by-task - Abstract Algebra | 32 | Accuracy (%) | 68.1 |
| MMLU-by-task - High School Physics | 32.45 | Accuracy (%) | 65.9 |
| MMLU-by-task - Global Facts | 35 | Accuracy (%) | 64.1 |
| ToolBench - WebShop Short | 0.7 | Task Score | 53.5 |
| ToolBench - WebShop Long | 0 | Task Score | 44.2 |
| MMLU-by-task - College Physics | 23.53 | Accuracy (%) | 41.5 |
| MMLU-by-task - Electrical Engineering | 37.93 | Accuracy (%) | 41.5 |
| MMLU-by-task - HellaSwag | 55.1 | Accuracy (%) | 40 |
| ToolBench - The Cat API | 73 | Task Score | 39.5 |
| MMLU-by-task - Human Sexuality | 39.69 | Accuracy (%) | 39.4 |
Interactive version: theaggregate.ai/model?slug=gpt-neoxt-chat-base-20b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.