What's New — daily benchmark digest
New benchmark leaders
- Mixedbread (+ Opus 5) took the lead on EnterpriseRAG Bench - Recall (94.28 vs 86.55 by Troml)
- Mixedbread (+ Opus 5) took the lead on EnterpriseRAG Bench (86.58 vs 80.34 by metor.com)
- Mixedbread (+ Opus 5) took the lead on EnterpriseRAG Bench - Correctness (89.8 vs 83.8 by Troml)
- LimiX-2 (default) took the lead on TabArena All Tasks (93.1 vs 87.3 by TabFM (default))
- Mixedbread (+ Opus 5) took the lead on EnterpriseRAG Bench - Completeness (90.62 vs 86.22 by metor.com)
- Muse-Glimmer-30B (high, 16k) took the lead on ReasonScape R12 (954.4 vs 951.64 by Qwen3.5-397B-A17B (AWQ, 16k) (Thinking))
- Mixedbread (+ Opus 5) took the lead on EnterpriseRAG Bench - Valid Extra Docs Resistance (99.61 vs 99.53 by OpenClaw)
New benchmarks
- Insurance Agent Benchmark (Cooper Harness)
- Insurance Agent Benchmark (Model Alone)
- Vals AI Vibe Code Bench 1-100
Interactive version: theaggregate.ai/whats-new · How It Works · Data refreshed daily, snapshot 2026-09-19.