CausalPhys - Entity Faithfulness: leaderboard

Metric: Entity Faithfulness (0-1 scaled to %), the share of causal-graph entities the rationale mentions, averaged over the four domains; the model writes a rationale and an answer; a GPT-4o judge makes binary checks of the rationale against the expert-annotated causal graph (typed object, attribute and event nodes with directed dependencies); higher is better. Source: arxiv.org. Saturation forecast: Around March 2028. 11 models tracked.

Top models

#ModelScore
1InternVL3-78B69.68
2GPT-4o67.95
3Qwen 3 VL 32B67.95
4GPT-4o Mini66.52
5Claude Sonnet 466.34
6Gemini 2.5 Flash65.97
7Mistral Small 3.262.87
8Phi-4 Multimodal Instruct62.26
9Qwen 2 VL 7B58.71

Interactive version: theaggregate.ai/benchmark?slug=causalphys-entity-faithfulness · How It Works · Data refreshed daily, snapshot 2026-09-29.