Unison - Mutual Enhancement: leaderboard
Metric: Unified score (%) on the mutual enhancement dimension of Unison (2,169 unified understanding-and-generation samples): multi-round tasks where understanding and generation correct each other, scored by the authors' Unison-Judge (a human-aligned 8B judge) and task metrics; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Pro | 71.4 |
| 2 | GPT-5.2 | 70.2 |
Interactive version: theaggregate.ai/benchmark?slug=unison-mutual-enhancement · How It Works · Data refreshed daily, snapshot 2026-09-29.