Bielik-11B-v2.0-Instruct — benchmark results
SpeakLeash and ACK Cyfronet AGH's Polish 11B instruct tune of Bielik-11B-v2, a depth-upscaled Mistral 7B v0.2 trained on curated Polish corpora (August 2024). Provider: SpeakLeash. Released 2024-08-26. Access: Open.
Unified ELO 1473 ± 12, rank #898 of 1776 rated models, from 22 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Open PL LLM - RAG | 75.5 | Average RAG Score (%) | 96.2 |
| Open PL LLM - Multiple Choice | 63.31 | Average Multiple-Choice Score (%) | 95.1 |
| Open PL LLM Leaderboard | 65.12 | Average Score (%) | 90.2 |
| Open PL LLM - Generative | 66.76 | Average Generative Score (%) | 83.2 |
| Open LLM Leaderboard - MuSR | 14.74 | Score | 80 |
| LLMZSZL Leaderboard | 55.61 | Score | 75.5 |
| Polish EQ-Bench | 68.24 | EQ-Bench Score | 75.2 |
| Open LLM Leaderboard - GPQA | 8.95 | Score | 73.2 |
| Open LLM Leaderboard - BBH | 33.77 | Score | 66.4 |
| MT-Bench PL - Writing | 8.75 | Judge Score (0-10) | 65.3 |
| MT-Bench PL - Coding | 5.6 | Judge Score (0-10) | 63.3 |
| Open LLM Leaderboard - IFEval | 52.52 | Score | 62.5 |
Interactive version: theaggregate.ai/model?slug=bielik-11b-v2-0-instruct · How the rankings work · Data refreshed daily, snapshot 2026-07-22.