Bielik-11B-v2.2-Instruct — benchmark results
SpeakLeash's Polish-focused 11B instruct fine-tune of the Bielik v2 base, a Mistral 7B v0.2 depth-upscaled to 11B (August 2024). Provider: SpeakLeash. Released 2024-08-26. Access: Open.
Unified ELO 1514 ± 17, rank #753 of 1776 rated models, from 23 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Open PL LLM - RAG | 76.92 | Average RAG Score (%) | 100 |
| Open PL LLM - Multiple Choice | 63.41 | Average Multiple-Choice Score (%) | 96.3 |
| Open PL LLM Leaderboard | 65.57 | Average Score (%) | 93 |
| MT-Bench PL - Writing | 9.35 | Judge Score (0-10) | 90.8 |
| Open PL LLM - Generative | 67.23 | Average Generative Score (%) | 85 |
| Open LLM Leaderboard - GPQA | 10.85 | Score | 81 |
| Open LLM Leaderboard - MATH Level 5 | 26.81 | Score | 79.4 |
| LLMZSZL Leaderboard | 57.36 | Score | 77.6 |
| MT-Bench PL - Roleplay | 9.03 | Judge Score (0-10) | 77.6 |
| Polish EQ-Bench | 69.05 | EQ-Bench Score | 77.2 |
| Open LLM Leaderboard - BBH | 36.96 | Score | 76.4 |
| MT-Bench PL - Reasoning | 6.9 | Judge Score (0-10) | 75.5 |
Interactive version: theaggregate.ai/model?slug=bielik-11b-v2-2-instruct · How the rankings work · Data refreshed daily, snapshot 2026-07-22.