Arcee Trinity Large (Thinking) — benchmark results
Arcee Trinity Large evaluated with thinking enabled. Provider: Arcee AI. Released 2026-04-01. Access: Open.
Unified ELO 1611 ± 20, rank #454 of 1776 rated models, from 26 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Position Bias (Lechmazur) | 36.6 | Order Flip % (lower is better) | 62.9 |
| Chatbot Arena (Text) | 1369 | Elo | 56.6 |
| PACT (Lechmazur) | 1526 | CMS Points | 56 |
| SpeechMap Compliance | 65.5 | % Requests Completed | 54.5 |
| BenchLM | 52.3 | Overall Score | 50.8 |
| Vals AI CaseLaw v2 | 57.88 | Accuracy (%) | 47.7 |
| PM-LLM-Benchmark | 30.2 | Score | 45.3 |
| Bullshit Benchmark | 20 | BS Detection Rate (%) | 42.2 |
| Design Arena (Website) | 1162 | Elo | 40.4 |
| Design Arena (3D) | 1139 | Elo | 39.8 |
| SnakeBench | 19.6 | TrueSkill Rating | 37.7 |
| Design Arena (Game Dev) | 1134 | Elo | 31.4 |
Interactive version: theaggregate.ai/model?slug=arcee-trinity-large-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.