AGC-Bench - grapheval_ai_researcher — leaderboard
Metric: Dataset z-score. Source: huggingface.co. 82 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GLM-5 | 1.88 |
| 2 | nemotron-3-super-120B-a12B | 1.42 |
| 3 | Gemma 4 26B A4B (IT) | 1.31 |
| 4 | nova-2-lite-v1 | 1.11 |
| 5 | Gemini 2.5 Flash | 1.08 |
| 6 | Claude Haiku 4.5 | 1.06 |
| 7 | ERNIE 4.5 300B A47B | 1.02 |
| 8 | GLM-4.7 | 0.91 |
| 9 | Mixtral 8x22B Instruct | 0.81 |
| 10 | DeepSeek V3.1 Terminus | 0.78 |
| 11 | GPT-5.5 | 0.73 |
| 12 | Claude Sonnet 4 | 0.72 |
| 13 | DeepSeek V3 Chat | 0.7 |
| 14 | Nova Pro (v1) | 0.67 |
| 15 | GPT-4.1 Mini | 0.67 |
Interactive version: theaggregate.ai/benchmark?slug=agc-bench-grapheval-ai-researcher · How the rankings work · Data refreshed daily, snapshot 2026-07-22.