Era by Eon - Hidden Knowledge: leaderboard

Metric: Correct runs (% of 24: the eight hidden-knowledge questions of this paper, each asked three times; a run is correct only if every field of its JSON answer equals the code-computed key, abstentions and runs that hit a limit count as wrong; read-only access to one generated large fintech company (one fixed seed; 250 accounts, 3,000 tickets, 3,506 recorded calls) through simulated Salesforce, HubSpot, Zendesk, Jira, Gong, Slack and S3 APIs; the Era agent program cannot run code (up to 165 turns and tool calls), the LangGraph program adds a Python sandbox (up to 150 turns, one hour); read as count x 100/24). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 12 models tracked.

Top models

#ModelScore
1Era Agent + Claude Fable 5.175
2LangGraph + Claude Fable 5.162.5
3Era Agent + GPT-6 Astra54.17
4LangGraph + GPT-6 Astra50
5Era Agent + Claude Sonnet 525
6LangGraph + Claude Sonnet 525
7LangGraph + GPT-5.6 Sol8.33
8Era Agent + GPT-5.6 Sol4.17
9Era Agent + GPT-5.6 Luna4.17
10LangGraph + GPT-5.6 Luna0

Interactive version: theaggregate.ai/benchmark?slug=era-by-eon-hidden-knowledge · How It Works · Data refreshed daily, snapshot 2026-09-26.