Open PL LLM - PolQA Reranking (multiple choice, 5-shot): leaderboard

Metric: Accuracy (%). Source: huggingface.co. 319 models tracked.

Top models

#ModelScore
1Bielik-11B-v2.1-Instruct85.63
2Bielik-11B-v2.2-Instruct85.4
3Bielik-11B-v2.3-Instruct84.66
4Qwen 2 72B84.6
5Qwen 2.5 32B Instruct84.37
6Qwen 2.5 32B83.53
7MSH-v1-Bielik-v2.3-Instruct-MedIT-merge83.34
8Qwen 2 72B Instruct83.15
9Llama 3.3 70B Instruct82.87
10Mistral Large 2 (Jul)82.78
11Qwen 3 14B82.57
12Bielik-11B-v2.0-Instruct82.45
13Llama 3.1 70B Instruct82.4
14Llama 3.1 Nemotron 70B Instruct82.37
15Qwen 2.5 14B Instruct82.19

Interactive version: theaggregate.ai/benchmark?slug=open-pl-llm-polqa-reranking-multiple-choice-5-shot · How It Works · Data refreshed daily, snapshot 2026-09-19.