AGC-Bench - mops — leaderboard

Metric: Dataset z-score. Source: huggingface.co. 83 models tracked.

Top models

#ModelScore
1Mistral Medium 3.12.56
2Llama 3 8B Instruct2.05
3Ministral 3 8B1.99
4Claude Sonnet 4.61.68
5nemotron-3-super-120B-a12B1.53
6Mistral Large 31.46
7Llama 3.1 70B Instruct1.32
8Gemma 3 4B (IT)1.08
9nova-2-lite-v11.05
10Claude Opus 4.61.02
11DeepSeek V3 (0324)1.01
12Claude Opus 4.71
13Kimi K2 09051
14GPT-5.10.96
15GPT-5.4 Nano0.81

Interactive version: theaggregate.ai/benchmark?slug=agc-bench-mops · How the rankings work · Data refreshed daily, snapshot 2026-07-22.