Trinity Large (Thinking): benchmark results
Trinity Large evaluated with thinking enabled. Provider: Arcee AI. Released 2026-04-01. Access: Open.
Unified ELO 1584 ± 1, rank #500 of 1761 rated models, from 38 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| CritPt | 90 | Accuracy (self-reported) | 97.7 |
| AA TAU-2 Bench | 90.06 | Accuracy (%) | 84.6 |
| AA Omniscience - Software Engineering (SWE) - R | 16 | Accuracy (%) | 72.5 |
| AA Omniscience - Software Engineering (SWE) - Java | 19 | Accuracy (%) | 71 |
| AA Omniscience - Software Engineering (SWE) - Swift | 40 | Accuracy (%) | 70.1 |
| AA IFBench | 56.26 | Accuracy (%) | 67.3 |
| AA Omniscience - Science, Engineering & Mathematics | 33.7 | Accuracy (%) | 67.2 |
| AA Humanity's Last Exam | 15.85 | Accuracy (%) | 66.8 |
| AA Omniscience - Software Engineering (SWE) - PHP | 26 | Accuracy (%) | 65 |
| AA Omniscience - Software Engineering (SWE) - Rust | 52 | Accuracy (%) | 64.7 |
| AA Terminal-Bench Hard | 22.73 | Accuracy (%) | 62.6 |
| AA CritPt | 0.86 | Accuracy (%) | 62.5 |
Interactive version: theaggregate.ai/model?slug=trinity-large-thinking · How It Works · Data refreshed daily, snapshot 2026-09-05.