Not-WizardLM-2-7B: benchmark results

Provider: Other. Access: Open.

Unified ELO 1455 ± 20, rank #1683 of 2928 rated models, from 11 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Open LLM Leaderboard v1 - TruthfulQA MC256.98MC2 (%) (0-shot)69.2
Open LLM Leaderboard v1 - ARC Challenge62.88Normalized accuracy (%) (25-shot)60.2
Open LLM Leaderboard v1 - GSM8K43.75Accuracy (%) (5-shot)58
Open LLM Leaderboard v1 - HellaSwag83.26Normalized accuracy (%) (10-shot)57.2
Open Portuguese LLM - OAB Exams43.14Accuracy (%)54.2
Open LLM Leaderboard v1 - MMLU61.53Accuracy (%) (5-shot)52.1
Open Portuguese LLM - BLUEX49.93Accuracy (%)46.2
Open Portuguese LLM - ENEM59.69Accuracy (%)46.1
Open Portuguese LLM - FaQuAD NLI61.55Macro F1 (%)46
Open Portuguese LLM - ASSIN2 RTE88.11Macro F1 (%)43.5
Open LLM Leaderboard v1 - WinoGrande73.56Accuracy (%) (5-shot)32.6

Interactive version: theaggregate.ai/model?slug=not-wizardlm-2-7b · How It Works · Data refreshed daily, snapshot 2026-09-23.