Skip to content
AI Atlas
PaperActive

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

arxiv.org/abs/2609.11897

Updated 30 min ago · first seen 11 Sept 2026

paper_01M294FPGV3HBX4DTBXWWTPHDP

Published
11 Sept 2026
T1 · 30 min ago
arXiv
2609.11897
T1 · 30 min ago
Category
cs.LG
T1 · 30 min ago

Abstract

Causal discovery aims to uncover causal structures from data and is fundamental to scientific reasoning and intervention-based decision making. Its evaluation relies heavily on structural causal models (SCMs), which specify a causal graph together with the mechanisms that generate data, yet existing studies differ substantially in graph families, mechanisms, and evaluation protocols. The emergence of causal discovery foundation models (CDFMs) further complicates evaluation: performance may reflect not only causal discovery ability, but also overlap between pretraining environments and test SCMs, making results on fixed synthetic benchmarks difficult to interpret. We introduce CausalArena, a unified and evolvable benchmark for causal discovery under a common protocol. Synthetic SCMs supply controlled breadth over structures and mechanisms; semantic operational SCMs provide human-auditable, semantically grounded environments beyond standard synthetic generators; and formula-grounded SCMs test discovery under explicit scientific mechanisms. Public real-world datasets provide an additional external-validity check. Experiments across classical, neural, and pretrained methods reveal substantial ranking shifts across SCM families and protocols, showing that strong performance in one benchmark regime does not reliably transfer to others. These results highlight benchmark diversity and pretraining--evaluation overlap as central challenges for evaluating causal discovery in the foundation model era.

Authors 4

Zi-Rong Li, Si-Yang Liu, Tian-Zuo Wang, Han-Jia Ye

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 30 min agohigh

Arxiv announce type
new

Source:arXiv (Atom API + RSS)T1observed 30 min agohigh

arXiv id
2609.11897

Source:arXiv (Atom API + RSS)T1observed 30 min agohigh

Categories
cs.LG

Source:arXiv (Atom API + RSS)T1observed 30 min agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 30 min agohigh

Primary category
cs.LG

Source:arXiv (Atom API + RSS)T1observed 30 min agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 30 min agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

30 min ago

Conflicts

None