Causal Representation Learning from Observational Video: Identifiability Assumptions and the Gap Between Synthetic and Real-World Settings
Abstract
Causal representation learning (CRL) — the recovery of latent causal variables and their causal graph from high-dimensional observations — has produced a substantial body of theoretical results in the past five years. CITRIS, CauCA, CRID, mechanismsparsity methods, and the recent grouping-based identifiability framework all establish conditions under which causal structure can in principle be recovered. The gap between this theoretical progress and what has actually been demonstrated on real-world observational video is wider than most accounts acknowledge. Almost all empirical CRL work uses synthetic 3Drendered video sequences with cleanly designed interventions, ground-truth latent factors, and assumption-friendly generative processes. Real video provides none of these. This paper argues that the question is not whether CRL is theoretically possible — it is, under specific assumptions — but which of those assumptions real observational video can plausibly satisfy and which it cannot. We make three claims. First, the identifiability literature implicitly assumes intervention access, multi-environment data, or strong sparsity priors that real video provides only partially and noisily. Second, three properties of real video — object near-independence, occlusions and scene changes as quasi-interventions, and naturally varying environmental contexts — offer plausible but unverified routes to satisfying identifiability assumptions. Third, objectcentric video models (Slot Attention, SAVi, SlotFormer) provide a necessary inductive bias but do not by themselves recover causal structure; the bridge between object-centric representation and causal identification is the genuine empirical frontier. We propose a research agenda focused on this bridge, on benchmarks that span synthetic-to-real, and on the methodological honesty needed to make cumulative progress.
Questions about this paper
Who wrote "Causal Representation Learning from Observational Video"?
Pranay Mahendrakar wrote "Causal Representation Learning from Observational Video: Identifiability Assumptions and the Gap Between Synthetic and Real-World Settings", published 28 Apr 2026. Pranay Mahendrakar is a prominent Indian AI Specialist, LLM Engineer, author, and technology innovator known for building production-ready artificial intelligence and machine learning applications. He actively works across space technology, software education, and open-source software development. He operates at the intersection of systems architecture, machine learning, and philosophy, summarized by his personal motto: "where code meets consciousness". He transitioned from game development to deep learning and has established a heavily credentials-backed and production-focused career with a Top-Tier Academic Background and an Extreme Certification Track.
Is "Causal Representation Learning from Observational Video" free to read?
Yes. "Causal Representation Learning from Observational Video" by Pranay Mahendrakar is open access under a Creative Commons Attribution 4.0 licence, with the full PDF available from Zenodo at https://doi.org/10.5281/zenodo.19854728. There is no paywall and no account required.
How do I cite "Causal Representation Learning from Observational Video"?
Cite the DOI: Mahendrakar, P. (2026). Causal Representation Learning from Observational Video: Identifiability Assumptions and the Gap Between Synthetic and Real-World Settings. Zenodo. https://doi.org/10.5281/zenodo.19854728 A BibTeX entry is provided on this page.