Era by Eon - Computable Questions (Large Company): leaderboard
Metric: Correct answers (% of the 27 computable questions of the current benchmark on the large company, one run each; every rule the answer needs is stated in the question or a company document; read-only access to one generated large fintech company (one fixed seed; 250 accounts, 3,000 tickets, 3,506 recorded calls) through simulated Salesforce, HubSpot, Zendesk, Jira, Gong, Slack and S3 APIs; the Era agent program cannot run code (up to 165 turns and tool calls), the LangGraph program adds a Python sandbox (up to 150 turns, one hour); read as count x 100/27). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 10 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | LangGraph + Claude Fable 5.1 | 92.59 |
| 2 | LangGraph + Claude Sonnet 5 | 92.59 |
| 3 | LangGraph + GPT-6 Astra | 88.89 |
| 4 | LangGraph + GPT-5.6 Sol | 81.48 |
| 5 | Era Agent + GPT-6 Astra | 29.63 |
| 6 | LangGraph + GPT-5.6 Luna | 29.63 |
| 7 | Era Agent + GPT-5.6 Sol | 3.7 |
| 8 | Era Agent + DeepSeek-V3.2 | 3.7 |
| 9 | Era Agent + GPT-5.6 Luna | 0 |
| 10 | LangGraph + DeepSeek-V3.2 | 0 |
Interactive version: theaggregate.ai/benchmark?slug=era-by-eon-computable-questions-large-company · How It Works · Data refreshed daily, snapshot 2026-09-26.