free-evo-qwen72B-v0.8-re — benchmark results
Freewheelin's evolutionary merge of Qwen2-based 72B models, inspired by Sakana AI's evolutionary model-merging method. Provider: Other. Released 2024-05-02. Access: Open.
Unified ELO 1575 ± 9, rank #546 of 1776 rated models, from 137 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Open Chinese LLM - ARC Challenge | 71.16 | Accuracy (%) | 99.9 |
| Open Chinese LLM - HellaSwag | 78.09 | Accuracy (%) | 99.4 |
| Open Chinese LLM Leaderboard | 74.62 | Average Score (%) | 99.4 |
| Open Arabic LLM - Arabic MMLU HT Global Facts | 54 | Accuracy (%) | 98.8 |
| Open Arabic LLM - Madinah QA Arabic Language (Grammar) | 70.96 | Accuracy (%) | 98.8 |
| Open Chinese LLM - CMMLU | 73.88 | Accuracy (%) | 98.8 |
| Open Chinese LLM - C-Eval Semantic | 91.21 | Accuracy (%) | 98.5 |
| Open Chinese LLM - TruthfulQA MC | 66.27 | Accuracy (%) | 97.6 |
| Open Arabic LLM - Alghafa Multiple Choice Rating Sentiment Task | 57.83 | Accuracy (%) | 97.5 |
| Open Arabic LLM - Arabic MMLU Accounting (University) | 74.32 | Accuracy (%) | 97.2 |
| Open Chinese LLM - GSM8K | 70.28 | Accuracy (%) | 96.4 |
| Open Chinese LLM - WinoGrande | 71.43 | Accuracy (%) | 96.4 |
Interactive version: theaggregate.ai/model?slug=free-evo-qwen72b-v0-8-re · How the rankings work · Data refreshed daily, snapshot 2026-07-22.