MAmmoTH2-8x7B-Plus — benchmark results
TIGER-Lab's (U. Waterloo) Mixtral 8x7B tune on 10M web-mined WebInstruct pairs, with extra public instruction data for reasoning and chat (May 2024). Provider: CMU. Released 2024-05-06. Access: Open.
Unified ELO 1450 ± 15, rank #1002 of 1776 rated models, from 10 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Science Leaderboard | 55.5 | Average Accuracy (%) | 90.9 |
| Open Chinese LLM - GSM8K | 56.56 | Accuracy (%) | 78.9 |
| Open Chinese LLM - TruthfulQA MC | 58.49 | Accuracy (%) | 77.2 |
| Open Chinese LLM - ARC Challenge | 57.42 | Accuracy (%) | 74.5 |
| Open Chinese LLM Leaderboard | 61.5 | Average Score (%) | 73 |
| MixEval | 51.8 | Score | 70.6 |
| Open Chinese LLM - C-Eval Semantic | 77.93 | Accuracy (%) | 66.9 |
| Open Chinese LLM - HellaSwag | 61.38 | Accuracy (%) | 62.8 |
| Open Chinese LLM - CMMLU | 56.53 | Accuracy (%) | 61.4 |
| Open Chinese LLM - WinoGrande | 62.19 | Accuracy (%) | 37.8 |
Interactive version: theaggregate.ai/model?slug=mammoth2-8x7b-plus · How the rankings work · Data refreshed daily, snapshot 2026-07-22.