OpenChat-3.5-0106-Gemma: benchmark results

Provider: Other. Access: Open.

Unified ELO 1468 ± 13, rank #1389 of 2656 rated models, from 54 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Open Portuguese LLM - FaQuAD NLI84.26Macro F1 (%)97.8
Open Portuguese LLM - ASSIN2 RTE93.99Macro F1 (%)93.8
Open PL LLM - PoQuAD Open Book (generative, 5-shot)70.28Levenshtein Similarity (%)85.9
Open PL LLM - DYK (generative, 5-shot)68.28Binary F1 (%)85.8
Open PL LLM - EQ-Bench55.55EQ-Bench Score84.6
Open PL LLM - PolQA Reranking (multiple choice, 5-shot)80.06Accuracy (%)83
Open PL LLM - DYK (multiple choice, 5-shot)67.35Binary F1 (%)82.8
Open Korean LLM - Ko-GSM8K (flexible extract)46.4Exact match, flexible extract (%)80.8
Open PL LLM - RAG70.28Average RAG Score (%)79
Open PL LLM - PPC (multiple choice, 5-shot)77.1Accuracy (%)75.4
Open PL LLM - PSC (multiple choice, 5-shot)81.65Binary F1 (%)74.3
Open PL LLM - PolEmo2-IN (generative, 5-shot)82.41Accuracy (%)74.2

Interactive version: theaggregate.ai/model?slug=openchat-3-5-0106-gemma · How It Works · Data refreshed daily, snapshot 2026-09-19.