openbuddy-zero-14B-v22.3-32k — benchmark results

OpenBuddy's multilingual 14B chat model with 32k context, assembled from Yi-1.5-9B-32K and DeepSeek-VL weights and further trained on 8 languages. Provider: Other. Released 2024-07-16. Access: Open.

Unified ELO 1407 ± 15, rank #1192 of 1776 rated models, from 14 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Open Chinese LLM - GSM8K50.11Accuracy (%)67.8
Open LLM Leaderboard - GPQA7.61Score64.1
Open Chinese LLM - ARC Challenge54.01Accuracy (%)61.9
Open Chinese LLM - TruthfulQA MC55.03Accuracy (%)60.5
Open LLM Leaderboard - MuSR11.34Score59.2
Open LLM Leaderboard - MATH Level 59.37Score45.3
Open LLM Leaderboard - MMLU-Pro24.3Score43.4
Open LLM Leaderboard - BBH26.29Score40.9
Open Chinese LLM Leaderboard55.41Average Score (%)39.8
Open LLM Leaderboard - IFEval37.53Score36.6
Open Chinese LLM - CMMLU50.45Accuracy (%)36.2
Open Chinese LLM - C-Eval Semantic64.01Accuracy (%)35

Interactive version: theaggregate.ai/model?slug=openbuddy-zero-14b-v22-3-32k · How the rankings work · Data refreshed daily, snapshot 2026-07-22.