论文

arXiv 论文证明类 Laplace 源下非线性 ICA 可精确识别

Kinks vs. Smoothness: Identifiability of Real Analytic nICA for Laplace-like Sources

精选理由

用 Laplace 分布的折点特性证明了非线性 ICA 能精确恢复隐因子,CelebA 实验里真拆出了控制单一属性的因子

一篇 arXiv 论文针对非线性独立成分分析(nICA)的核心难题给出证明:如何从混合数据中精确恢复真实的隐藏独立因子。论文证明,当源概率密度函数的一阶导数仅有有限个不连续点时(Laplace 分布是最典型的例子),实解析生成函数具有可识别性,即最多存在平凡歧义。证明思路是利用源分布中的折点(kinks)与实解析函数平滑性之间的对比。由于 tanh、softplus、GELU 等标准激活函数可逼近实解析函数,该结论可直接套用在 Normalizing Flows 和 Variational Autoencoders 的现有训练流程上。在 CelebA 数据实验中,作者恢复了多个控制单一属性的可解释潜在因子。

原文 · arXiv cs.LG

Kinks vs. Smoothness: Identifiability of Real Analytic nICA for Laplace-like Sources

Many machine learning systems try to explain complex data - like images or financial time series - in terms of hidden, independent factors that generated them. Recovering the true underlying factors, rather than some scrambled version of them, is the central challenge of nonlinear Independent Component Analysis (nICA). We prove identifiability (exact recovery) up to trivial ambiguities for real analytic generating functions when source probability density functions have a finite number of discontinuities in the first derivative. The Laplace distribution is the most prominent example satisfying this assumption. Our proof relies on the contrast between kinks in the source distribution and the smoothness of real analytic functions. Real analytic functions comprise a broad class of generating mechanisms, and can be approximated with Normalizing Flows or Variational Autoencoders with standard activation functions (e.g., tanh, softplus, GELU), so our result applies with minimal changes to existing training pipelines. We perform experiments on real and synthetic data with both Normalizing Flows and Variational Auto-Encoders demonstrating their identifiability properties. In experiments on CelebA data we recover several interpretable latent factors controlling unique attributes across the dataset.