Skip to content
AI Atlas

Sanity Checking Causal Representation Learning on a Simple Real-World System

Published 16 Sept 2026arXiv:2502.20099

data quality89

Updated 11 h ago · first seen 15 Sept 2026

paper_01M2JK0D23YFDF8GKHEMRVCM7W

Abstract

We evaluate methods for causal representation learning (CRL) on a simple, real-world system that satisfies the basic problem setup of CRL. The system consists of a controlled optical experiment producing a variety of measurements where the underlying causal factors---the control inputs to the experiment---are known, providing a ground truth. We select methods representative of different approaches to CRL and find that they all fail to consistently recover the underlying causal factors. To understand the failure modes of the evaluated algorithms, we perform an ablation on the data by substituting the real data-generating process with a simpler synthetic equivalent. The results reveal a reproducibility problem, as most methods already fail on this synthetic ablation despite its simple data-generating process. Additionally, we observe that common assumptions on the mixing function are crucial for the performance of some of the methods but do not hold in the real data. Our efforts highlight the contrast between the theoretical promise of the state of the art and the challenges in its application. We hope the benchmark serves as a simple, real-world sanity check to further develop and validate methodol- ogy, bridging the gap towards CRL methods that work in practice.

Authors

Authors 3

Jakob RungeJuan L. GamellaSimon Bing

Linked names open researcher pages (created from the paper's author list; name-only, no affiliation unless a source states it). Unlinked names have no researcher record yet.

Organizations

Organizations 0

No organization stated. arXiv metadata does not carry affiliations; an organization is linked only when a model card or lab page cites the paper.

Models

Models introduced or described 0

Inbound described_by relations from model cards and documentation.

No model links this paper yet

Model pages link papers through their model cards and documentation; the relation is written only when a source states it.

Datasets

Datasets used 0

No dataset relation recorded.

Benchmarks

Benchmarks used 0

No benchmark relation recorded.

Code

Repositories & frameworks 0

No repository linked.

Timeline

Timeline 2

Full timeline →

Sources

Sources 2

Source documents
SourceDocumentTypeTierLast observedSnapshots
arXiv (Atom API + RSS)rss.arxiv.org/rss/cs.AI feedT1· Official6 h ago5
arXiv (Atom API + RSS)rss.arxiv.org/rss/cs.LG feedT1· Official6 h ago4

Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.