Bielik-11B-v2.3-Instruct — benchmark results

SpeakLeash's Polish 11B instruct model, a linear merge of Bielik v2.0/2.1/2.2 built on a depth-upscaled Mistral 7B v0.2 (September 2024). Provider: SpeakLeash. Released 2024-09-05. Access: Open.

Unified ELO 1516 ± 11, rank #743 of 1776 rated models, from 298 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
EuroEval Bosnian Summarization - LR SUM BS32.58Score (%)100
EuroEval French Summarization - Orange SUM39.96Score (%)100
EuroEval Serbian Summarization - LR SUM SR31.95Score (%)100
EuroEval Spanish NLU - MLQA ES72.14Reading comprehension Score (%)100
EuroEval French NLU - Allocine97.49Sentiment classification Score (%)99.5
EuroEval Portuguese NLU - MultiWikiQA PT80.57Reading comprehension Score (%)99.5
EuroEval Romanian Summarization - Sumo RO37.99Score (%)99.3
EuroEval Spanish Summarization - Mlsum ES31.19Score (%)99.3
EuroEval Bulgarian NLU - Cinexio61.38Sentiment classification Score (%)99.1
EuroEval Romanian NLU - MultiWikiQA RO76.93Reading comprehension Score (%)99
Open PL LLM - RAG76.2Average RAG Score (%)98.3
EuroEval Bosnian58.74Average Score (%)97.1

Interactive version: theaggregate.ai/model?slug=bielik-11b-v2-3-instruct · How the rankings work · Data refreshed daily, snapshot 2026-07-22.