Trinity Large (Thinking) — benchmark results

Trinity Large evaluated with thinking enabled. Provider: Arcee AI. Released 2026-04-01. Access: Open.

Unified ELO 1671 ± 26, rank #319 of 1776 rated models, from 39 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AA TAU-2 Bench90.06Accuracy (%)84.6
CritPt0.9Accuracy (self-reported)74.5
AA Humanity's Last Exam14.69Accuracy (%)72.7
AA Omniscience - Science, Engineering & Mathematics33.3Accuracy (%)72.2
AA CritPt0.86Accuracy (%)68.8
AA IFBench56.26Accuracy (%)67.2
AA Omniscience - Health22.4Accuracy (%)64
AA Omniscience - Humanities & Social Sciences23.2Accuracy (%)63.1
AA Terminal-Bench Hard22.73Accuracy (%)62.6
AA GPQA Diamond75.15Accuracy (%)62.3
AA Omniscience - Law14Accuracy (%)61.8
AA-Omniscience Accuracy22.75Accuracy (%)61.4

Interactive version: theaggregate.ai/model?slug=trinity-large-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.