Bielik-11B-v2.3-Instruct — benchmark results
SpeakLeash's Polish 11B instruct model, a linear merge of Bielik v2.0/2.1/2.2 built on a depth-upscaled Mistral 7B v0.2 (September 2024). Provider: SpeakLeash. Released 2024-09-05. Access: Open.
Unified ELO 1516 ± 11, rank #743 of 1776 rated models, from 298 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| EuroEval Bosnian Summarization - LR SUM BS | 32.58 | Score (%) | 100 |
| EuroEval French Summarization - Orange SUM | 39.96 | Score (%) | 100 |
| EuroEval Serbian Summarization - LR SUM SR | 31.95 | Score (%) | 100 |
| EuroEval Spanish NLU - MLQA ES | 72.14 | Reading comprehension Score (%) | 100 |
| EuroEval French NLU - Allocine | 97.49 | Sentiment classification Score (%) | 99.5 |
| EuroEval Portuguese NLU - MultiWikiQA PT | 80.57 | Reading comprehension Score (%) | 99.5 |
| EuroEval Romanian Summarization - Sumo RO | 37.99 | Score (%) | 99.3 |
| EuroEval Spanish Summarization - Mlsum ES | 31.19 | Score (%) | 99.3 |
| EuroEval Bulgarian NLU - Cinexio | 61.38 | Sentiment classification Score (%) | 99.1 |
| EuroEval Romanian NLU - MultiWikiQA RO | 76.93 | Reading comprehension Score (%) | 99 |
| Open PL LLM - RAG | 76.2 | Average RAG Score (%) | 98.3 |
| EuroEval Bosnian | 58.74 | Average Score (%) | 97.1 |
Interactive version: theaggregate.ai/model?slug=bielik-11b-v2-3-instruct · How the rankings work · Data refreshed daily, snapshot 2026-07-22.