Bielik-11B-v2.3-Instruct: benchmark results

SpeakLeash's Polish 11B instruct model, a linear merge of Bielik v2.0/2.1/2.2 built on a depth-upscaled Mistral 7B v0.2 (September 2024). Provider: SpeakLeash. Released 2024-09-05. Access: Open.

Unified ELO 1564 ± 12, rank #775 of 2656 rated models, from 378 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
EuroEval Bosnian Summarization - LR SUM BS32.58Score (%)100
EuroEval French Summarization - Orange SUM39.96Score (%)100
EuroEval Serbian Summarization - LR SUM SR31.95Score (%)100
EuroEval Spanish NLU - MLQA ES72.14Reading comprehension Score (%)100
EuroEval French NLU - Allocine97.49Sentiment classification Score (%)99.5
EuroEval Portuguese NLU - MultiWikiQA PT80.57Reading comprehension Score (%)99.5
EuroEval Romanian Summarization - Sumo RO37.99Score (%)99.3
EuroEval Spanish Summarization - Mlsum ES31.19Score (%)99.3
EuroEval Bulgarian NLU - Cinexio61.38Sentiment classification Score (%)99.1
EuroEval Romanian NLU - MultiWikiQA RO76.93Reading comprehension Score (%)99
Open PL LLM - RAG76.2Average RAG Score (%)98.3
Open PL LLM - PolQA Open Book (generative, 5-shot)92.19Levenshtein Similarity (%)97.8

Interactive version: theaggregate.ai/model?slug=bielik-11b-v2-3-instruct · How It Works · Data refreshed daily, snapshot 2026-09-19.