What's New — daily benchmark digest
New benchmark leaders
- Gemini 3.1 Pro (Preview) took the lead on MMOU (82.67 vs 64.2 by Gemini 2.5 Pro)
- minimax-h3 took the lead on Design Arena (Video Editing) (1395 vs 1381 by Gemini 2.0 Flash)
- Claude Opus 5 (Claude Code) xHigh took the lead on SEAL - SWE Atlas - Test Writing (62.22 vs 55.6 by Claude Fable 5 (Claude Code) xHigh)
- Ovis-Omni-Embedding-v0.1-3B took the lead on MMEB Leaderboard (52.55 vs 46.53 by e5-omni-7B)
- Claude Opus 5 (Claude Code) xHigh took the lead on SEAL - SWE Atlas - Codebase QnA (63.17 vs 57.26 by Claude Opus 4.8 (Claude Code) xhigh)
- Claude Fable 5 took the lead on BridgeBench Security (1005.3 vs 1005.1 by GPT-5.6 Sol)
- Claude Fable 5 took the lead on BridgeBench Debugging (1005.3 vs 1005.1 by GPT-5.6 Sol)
- Claude Fable 5 took the lead on BridgeBench Refactoring (1005.3 vs 1005.1 by GPT-5.6 Sol)
- Claude Fable 5 took the lead on BridgeBench Hallucination (1005.3 vs 1005.1 by GPT-5.6 Sol)
New benchmarks
Interactive version: theaggregate.ai/whats-new · How It Works · Data refreshed daily, snapshot 2026-08-05.