Multivac Evaluation — leaderboard
Open peer-review evaluation where LLMs rate each other's responses. Aggregates peer scores across multiple evaluation rounds to measure general quality.
Metric: Avg Peer Score (0-10). Source: github.com. 19 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Grok 4.1 Fast | 9.18 |
| 2 | Grok Code Fast 1 | 9.12 |
| 3 | MiMo-V2-Flash | 8.96 |
| 4 | Claude Opus 4.5 | 8.95 |
| 5 | DeepSeek V3 | 8.93 |
| 6 | Claude Sonnet 4.5 | 8.92 |
| 7 | Gemini 3 Flash | 8.92 |
| 8 | Gemini 2.5 Flash | 8.72 |
| 9 | Gemini 2.5 Flash Lite | 8.72 |
| 10 | GPT-OSS-120B | 8.58 |
| 11 | GLM-4.7 | 7.28 |
| 12 | Gemini 3 Pro | 7.2 |
| 13 | MiniMax-M2 | 5.93 |
Interactive version: theaggregate.ai/benchmark?slug=multivac-evaluation · How the rankings work · Data refreshed daily, snapshot 2026-07-22.