TemporalBench — leaderboard

TemporalBench evaluates LLM-based agents on contextual and event-informed time-series tasks spanning multiple datasets and task types.

Metric: Overall MCQ Accuracy (%). Source: huggingface.co. 5 models tracked.

Top models

#ModelScore
1Single LLM (gpt-4o)34.11
2AgentScope (gpt-4o)32.91
3CAMEL (gpt-4o)32.33
4MetaGPT (gpt-4o)31.56
5TimeSeries Scientist (gpt-4o)23.96

Interactive version: theaggregate.ai/benchmark?slug=temporalbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.