Clembench Multimodal v1.6.5: leaderboard
Three image-based dialogue games (reference game, MatchIt, MapWorld navigation) from the University of Potsdam clembench (2024); clemscore = share of games completed times mean quality score.
Metric: clemscore (self-reported). Source: benchmarklist.com. Status: saturated. 20 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude 3.5 Sonnet (20240620) | 80.77 |
| 2 | GPT-4o (2024-08-06) | 80.04 |
| 3 | GPT-4 Preview (1106) | 73.55 |
| 4 | GPT-4o (2024-05-13) | 69.56 |
| 5 | Claude 3 Opus (20240229) | 68.16 |
| 6 | Gemma 3 27B (IT) | 61.39 |
| 7 | GPT-4o Mini (2024-07-18) | 58.46 |
| 8 | Gemini 1.5 Flash | 47.73 |
| 9 | Pixtral-12B-2409 | 28.64 |
| 10 | InternVL2-8B | 23.17 |
Interactive version: theaggregate.ai/benchmark?slug=clembench-multimodal-v1-6-5 · How It Works · Data refreshed daily, snapshot 2026-09-05.