K-EXAONE (Thinking): benchmark results

LG AI Research's open-weight K-EXAONE-236B-A23B (December 2025), a 236-billion-parameter, 23-billion-active MoE model covering six languages with a 256K context and Multi-Token Prediction for faster decoding. Provider: LG AI. Released 2025-12-31. Access: Open.

Unified ELO 1555 ± 1, rank #637 of 2032 rated models, from 108 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Nejumi 4 - Toxicity - Fairness99.44Criteria met (%)93.8
Nejumi 4 - Toxicity - Social Norms98.81Criteria met (%)90.8
Horangi 4 - Ko-HalluLens (Nonexistent Entities)94Refusal rate (%)89.9
Korean CSAT 2026 (Easy Mode) - Mathematics100Points (out of 100)88.4
Horangi 4 - HRM8K95Accuracy (%)85.6
Horangi 4 - HAE-RAE Bench (Reading Comprehension)90Accuracy (%)77.4
Nejumi 4 - BFCL - Live AST72.22Accuracy (%)76.8
Nejumi 4 - HalluLens96Hallucination resistance (%)75.4
Horangi 4 - IFEval-Ko88Instruction-following accuracy (%)73.1
Nejumi 4 - jaster (2-shot) - JSICK82Exact match (%)72.1
Horangi 4 - GLP - Mathematical Reasoning92.5Score (%)70.2
Horangi 4 - Ko-TruthfulQA85Accuracy (%)67.8

Interactive version: theaggregate.ai/model?slug=k-exaone-thinking · How It Works · Data refreshed daily, snapshot 2026-09-26.