Causal Representation Learning from Observational Video: Identifiability Assumptions and the Gap Between Synthetic and Real-World Settings
Abstract
Causal representation learning (CRL) — the recovery of latent causal variables and their causal graph from high-dimensional observations — has produced a substantial body of theoretical results in the past five years. CITRIS, CauCA, CRID, mechanismsparsity methods, and the recent grouping-based identifiability framework all establish conditions under which causal structure can in principle be recovered. The gap between this theoretical progress and what has actually been demonstrated on real-world observational video is wider than most accounts acknowledge. Almost all empirical CRL work uses synthetic 3Drendered video sequences with cleanly designed interventions, ground-truth latent factors, and assumption-friendly generative processes. Real video provides none of these. This paper argues that the question is not whether CRL is theoretically possible — it is, under specific assumptions — but which of those assumptions real observational video can plausibly satisfy and which it cannot. We make three claims. First, the identifiability literature implicitly assumes intervention access, multi-environment data, or strong sparsity priors that real video provides only partially and noisily. Second, three properties of real video — object near-independence, occlusions and scene changes as quasi-interventions, and naturally varying environmental contexts — offer plausible but unverified routes to satisfying identifiability assumptions. Third, objectcentric video models (Slot Attention, SAVi, SlotFormer) provide a necessary inductive bias but do not by themselves recover causal structure; the bridge between object-centric representation and causal identification is the genuine empirical frontier. We propose a research agenda focused on this bridge, on benchmarks that span synthetic-to-real, and on the methodological honesty needed to make cumulative progress.