MAC-Bench (AutoGen): leaderboard
Metric: Compliance-weighted success rate (%): success rate times compliance rate (CSR, alpha 1); 4,128 scenarios synthesized from 847 atomic legal and security rules (GDPR, PIPL, EU AI Act, CWE, OWASP, CIS) in sandboxed environments, under combined high pressure (authority plus urgency), models orchestrated by the AutoGen multi-agent framework with a shared compliance charter, temperature 0.7; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 12 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | AutoGen + Claude-3 (MAC-Bench checkpoint unspecified) | 47.6 |
| 2 | AutoGen + Claude-3.5 (MAC-Bench checkpoint unspecified) | 43.1 |
| 3 | AutoGen + GPT-4o | 37.3 |
| 4 | AutoGen + GPT-5 | 34.6 |
| 5 | AutoGen + Gemini-2.5 (MAC-Bench checkpoint unspecified) | 30.1 |
| 6 | AutoGen + Gemini-3 (MAC-Bench checkpoint unspecified) | 27.6 |
| 7 | AutoGen + Qwen-3-32B | 26.6 |
| 8 | AutoGen + Qwen-3-8B | 22.7 |
| 9 | AutoGen + Llama-3.1-70B | 19.5 |
| 10 | AutoGen + DeepSeek-V3 | 18.6 |
Interactive version: theaggregate.ai/benchmark?slug=mac-bench-autogen · How It Works · Data refreshed daily, snapshot 2026-09-29.