What's New — daily benchmark digest
New benchmark leaders
- Gemini 3.6 Flash took the lead on RuneBench (7454 vs 7400 by GPT-5.6 Sol)
- Claude Fable 5 took the lead on HiL-Bench (56.33 vs 29.1 by GPT-5.5)
- Troml took the lead on EnterpriseRAG Bench - Completeness (81.84 vs 72.86 by OpenClaw)
- Troml took the lead on EnterpriseRAG Bench (76.79 vs 68.22 by OpenClaw)
- Troml took the lead on EnterpriseRAG Bench - Recall (86.55 vs 79.02 by OpenClaw)
- Grok 4.5 took the lead on Vals AI SkillsBench (66.03 vs 62.55 by GPT-5.5 Codex)
- Troml took the lead on EnterpriseRAG Bench - Correctness (83.8 vs 81.6 by OpenClaw)
- Grok 4.5 (xHigh) took the lead on BoxPwnr CTF Bench (56.73 vs 55.47 by GLM-5.1)
- Gemini 3.5 Flash Cyber took the lead on LLM Stats (CyberGym) (83.2 vs 83.1 by Claude Mythos Preview)
New benchmarks
- LLM Arena RU
- AA-Briefcase
- FutureSim
- ProfBench
- SalesBench
- DrugDiscoveryBench
- TextQuests (No Clues)
- TextQuests (With Clues)
- RevengeBench
- GameCraft-Bench
- Kagi LLM Benchmark
- SpeechMap Compliance
- SnitchBench
- MLS-Bench Lite
New models
Interactive version: theaggregate.ai/digest · How the rankings work · Data refreshed daily, snapshot 2026-07-22.