Bielik-11B-v2.0-Instruct — benchmark results

SpeakLeash and ACK Cyfronet AGH's Polish 11B instruct tune of Bielik-11B-v2, a depth-upscaled Mistral 7B v0.2 trained on curated Polish corpora (August 2024). Provider: SpeakLeash. Released 2024-08-26. Access: Open.

Unified ELO 1473 ± 12, rank #898 of 1776 rated models, from 22 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Open PL LLM - RAG75.5Average RAG Score (%)96.2
Open PL LLM - Multiple Choice63.31Average Multiple-Choice Score (%)95.1
Open PL LLM Leaderboard65.12Average Score (%)90.2
Open PL LLM - Generative66.76Average Generative Score (%)83.2
Open LLM Leaderboard - MuSR14.74Score80
LLMZSZL Leaderboard55.61Score75.5
Polish EQ-Bench68.24EQ-Bench Score75.2
Open LLM Leaderboard - GPQA8.95Score73.2
Open LLM Leaderboard - BBH33.77Score66.4
MT-Bench PL - Writing8.75Judge Score (0-10)65.3
MT-Bench PL - Coding5.6Judge Score (0-10)63.3
Open LLM Leaderboard - IFEval52.52Score62.5

Interactive version: theaggregate.ai/model?slug=bielik-11b-v2-0-instruct · How the rankings work · Data refreshed daily, snapshot 2026-07-22.