Composer 2.5: benchmark results
Provider: Other. Released 2026-05-18. Access: API.
Unified ELO 1712 ± 1, rank #41 of 1392 rated models, from 9 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| BenchmarkList ECI | 134.11 | Capability Index (ECI) | 82.1 |
| Vals AI SWE-bench Verified | 79.6 | Resolved (%) | 71.3 |
| Agent Security League - Functional Correctness | 75.4 | Functional Correctness (%) | 63.9 |
| Vals AI Vibe Code Bench | 49.61 | Accuracy (%) | 59.3 |
| Vals AI Terminal-Bench 2.1 | 58.43 | Accuracy (%) | 48.4 |
| Agents' Last Exam | 20.4 | Pass Rate (%) | 45.8 |
| Agent Security League - Security Correctness | 14 | Security Correctness (%) | 38.9 |
| CursorBench 3.1 | 56.1 | Score (%) | 37.7 |
| FrontierSWE | 34 | Dominance (%) | 34.4 |
Interactive version: theaggregate.ai/model?slug=composer-2-5 · How It Works · Data refreshed daily, snapshot 2026-09-05.