Era by Eon - Computable Questions (Large Company): leaderboard

Metric: Correct answers (% of the 27 computable questions of the current benchmark on the large company, one run each; every rule the answer needs is stated in the question or a company document; read-only access to one generated large fintech company (one fixed seed; 250 accounts, 3,000 tickets, 3,506 recorded calls) through simulated Salesforce, HubSpot, Zendesk, Jira, Gong, Slack and S3 APIs; the Era agent program cannot run code (up to 165 turns and tool calls), the LangGraph program adds a Python sandbox (up to 150 turns, one hour); read as count x 100/27). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 10 models tracked.

Top models

#ModelScore
1LangGraph + Claude Fable 5.192.59
2LangGraph + Claude Sonnet 592.59
3LangGraph + GPT-6 Astra88.89
4LangGraph + GPT-5.6 Sol81.48
5Era Agent + GPT-6 Astra29.63
6LangGraph + GPT-5.6 Luna29.63
7Era Agent + GPT-5.6 Sol3.7
8Era Agent + DeepSeek-V3.23.7
9Era Agent + GPT-5.6 Luna0
10LangGraph + DeepSeek-V3.20

Interactive version: theaggregate.ai/benchmark?slug=era-by-eon-computable-questions-large-company · How It Works · Data refreshed daily, snapshot 2026-09-26.