Era by Eon - Hidden Knowledge: leaderboard
Metric: Correct runs (% of 24: the eight hidden-knowledge questions of this paper, each asked three times; a run is correct only if every field of its JSON answer equals the code-computed key, abstentions and runs that hit a limit count as wrong; read-only access to one generated large fintech company (one fixed seed; 250 accounts, 3,000 tickets, 3,506 recorded calls) through simulated Salesforce, HubSpot, Zendesk, Jira, Gong, Slack and S3 APIs; the Era agent program cannot run code (up to 165 turns and tool calls), the LangGraph program adds a Python sandbox (up to 150 turns, one hour); read as count x 100/24). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 12 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Era Agent + Claude Fable 5.1 | 75 |
| 2 | LangGraph + Claude Fable 5.1 | 62.5 |
| 3 | Era Agent + GPT-6 Astra | 54.17 |
| 4 | LangGraph + GPT-6 Astra | 50 |
| 5 | Era Agent + Claude Sonnet 5 | 25 |
| 6 | LangGraph + Claude Sonnet 5 | 25 |
| 7 | LangGraph + GPT-5.6 Sol | 8.33 |
| 8 | Era Agent + GPT-5.6 Sol | 4.17 |
| 9 | Era Agent + GPT-5.6 Luna | 4.17 |
| 10 | LangGraph + GPT-5.6 Luna | 0 |
Interactive version: theaggregate.ai/benchmark?slug=era-by-eon-hidden-knowledge · How It Works · Data refreshed daily, snapshot 2026-09-26.