Foundations of Independent Component Analysis
Abstract: We present the mathematical foundations of linear independent component analysis (ICA) models based on standard literature in a self-contained note. It is aimed at readers with a background in measure-theoretic probability theory. We first develop the theory of the characteristic functions of probability measures on R<sup>d, including their analyticity and the way in which they determine and characterise the distributions. We then focus on several identifiability results of ICA models with successively strengthened assumptions on the sources: from merely non-constant, to non-Gaussian, to Gaussian-free independent sources. Under the strictest assumptions, we show that the independent sources are identifiable up to translation, permutation, scales and signs, and this even in the presence of additive Gaussian noise. Furthermore, we present the online equivariant gradient descent ICA algorithm for recovering the independent sources from data, in the standard complete noiseless non-Gaussian ICA setting.
- Provable ICA with Unknown Gaussian Noise, and Implications for Gaussian Mixtures and Autoencoders (2012)
- On the Identifiability of Sparse ICA without Assuming Non-Gaussianity (2024)
- Finite Rank Perturbations of Toeplitz Products on the Bergman Space (2020)
- The level of distribution of the sum-of-digits function of linear recurrence number systems (2019)
- Radii of convexity of integral operators (2018)
- Corona problem with data in ideal spaces of sequences (2015)
- Identifiability and Estimation in High-Dimensional Nonparametric Latent Structure Models (2025)
- Diverse Influence Component Analysis: A Geometric Approach to Nonlinear Mixture Identifiability (2025)
- Finite-Sample Analysis of Nonlinear Independent Component Analysis:Sample Complexity and Identifiability Bounds (2026)
- Sinkhorn Linearization and the Spectral Proxy: Unifying the Statistical and Algorithmic Theory of Feature-Parameterized Inverse Optimal Transport via a Single Spectral Sandwich (2026)
Summary
- The paper unifies linear ICA identifiability through the Kagan–Linnik–Rao theorem, showing that non-Gaussian sources are recoverable up to scale, permutation, and translation, while Gaussian sources create fundamental ambiguities.
- Gaussian-free source decomposition extends identifiability to additive Gaussian noise with arbitrary, dependent, or degenerate covariance, enabling recovery of both source distributions and noise laws under two-sided assumptions.
- The paper connects equivariant natural-gradient stability to identifiability, showing that Gaussian sources lie on the instability boundary and that kurtosis is reliable only for cubic scores, not as a general measure of non-Gaussianity.
Scope and structure of the note
This self-contained monograph-style paper develops the mathematical foundations of linear independent component analysis (ICA) for readers with a measure-theoretic probability background. The model throughout is X=AZ+μ, where A∈Rp×k is an unobserved mixing matrix and Z has mutually independent components (the sources). The central question is identifiability: how much of the pair (A,L(Z)) is determined by the law of X alone. The paper's logical skeleton rests on a single engine — the Kagan–Linnik–Rao theorem on representations with non-constant independent sources — from which two branches grow: one strengthening the hypotheses on the sources to buy back stronger conclusions, the other fixing the complete noiseless case and analyzing estimation. Notably, every result is proved in the paper itself with a single exception (Bochner's theorem), including full proofs of Cramér's decomposition theorem, Marcinkiewicz' theorem, Lévy's continuity theorem, and the Kagan–Linnik–Rao theorem via a finite-difference argument.
Characteristic functions as the working language
All arguments are conducted at the level of characteristic functions. Three elementary identities carry the entire ICA theory: affine images transform characteristic functions by composition and multiplication by a phase; independent sums factor; and one-dimensional projections evaluate the characteristic function along directions. The paper develops cumulants through the distinguished logarithm, proves analyticity in a strip under exponential moment conditions, and derives a rigidity corollary: if two laws have exponential moments and their characteristic functions agree merely near the origin, then the laws agree globally. This matters because agreement on a neighbourhood of the origin does not suffice in general — the tent function (1−∣t∣)+ and its periodic extension coincide on [−1,1] yet are characteristic functions of an absolutely continuous law and a purely atomic law respectively.
Two classical characterisation results anchor everything downstream. Marcinkiewicz' theorem: if a characteristic function equals exp(g) near the origin with g a polynomial, then deg(g)≤2, i.e., the law is (possibly degenerate) Gaussian. Cramér's decomposition theorem: a Gaussian sum of independent variables has only Gaussian factors. The proof of Marcinkiewicz' theorem given here is a clean ridge-function argument using the Phragmén–Lindelöf-type "ridge property" that A∈Rp×k0 is maximized on each horizontal line at its imaginary-axis point; the proof of Cramér's theorem replaces Hadamard factorization with a Borel–Carathéodory coefficient bound.
The paper also devotes careful attention to kurtosis, and makes claims that contradict common practice. It shows that excess kurtosis measures dispersion about A∈Rp×k1 (Moors' reading), not tail weight or peakedness: it constructs a three-point law with compact support but arbitrarily large positive kurtosis, and a uniform-plus-Laplace mixture (A∈Rp×k2) whose density exceeds the same-variance Gaussian density beyond A∈Rp×k3 by a factor of roughly A∈Rp×k4 at A∈Rp×k5, while having negative kurtosis (A∈Rp×k6). It also flags that "sub-Gaussian" in the ICA sense (sign of kurtosis) is logically unrelated to "sub-Gaussian" in the concentration sense, and that mesokurtosis does not imply normality (the three-point law with A∈Rp×k7). The practical upshot: the correct hypothesis for ICA identifiability is non-normality, not non-zero kurtosis.
Identifiability with non-constant sources
The foundational result (proved in full in an appendix) compares two normalized representations of the same random vector, requiring only that sources be mutually independent and almost surely non-constant. Its first part is elementary: the affine hull of the support determines A∈Rp×k8 and the offsets agree modulo that image, so ranks are equal. Its second part is a column dichotomy: every column of A∈Rp×k9 either is proportional to a column of Z0 — in which case the corresponding source characteristic functions satisfy
Z1
near the origin for some polynomial Z2 — or is not, in which case the source Z3 is forced to be Gaussian. The proof is a finite-difference argument: taking distinguished logarithms yields a linear relation among ridge functions; difference operators in orthogonal directions annihilate all but one ridge term; Fréchet's functional equation (proved here in local form via a smoothing bootstrap) forces the survivor to be a polynomial; and Marcinkiewicz collapses the resulting degree bound to 2. The paper notes that this theorem generalizes the Darmois–Skitovich theorem (which is the case Z4), and it reinstates a proportionality constant missing from the statement in Kagan–Linnik–Rao's monograph.
Non-Gaussian sources and additive Gaussian noise
Specializing to non-Gaussian sources, the key technical device is a "noise-trade" lemma showing how Gaussian noise can be moved between the additive noise vector and extra columns of the mixing matrix, using a geometric lemma that rotates any frame away from finitely many forbidden directions. The main theorem establishes that two representations with independent non-Gaussian sources and Gaussian noise vectors (possibly degenerate, arbitrarily dependent across coordinates) must have mixing matrices related by Z5 with Z6 a permutation and Z7 invertible diagonal. When Z8 has full column rank, each source pair satisfies one of two distributional identities: either Z9 or the reverse, with (A,L(Z))0 Gaussian. Thus four ambiguities remain: translation, scale, permutation, and componentwise additive Gaussian noise. The first three are removable by convention; the fourth changes the source law itself and motivates the next section.
Gaussian-free sources: the strongest identifiability result
The paper introduces the notion of a Gaussian-free random variable: one admitting no non-degenerate Gaussian convolution factor. Quantitatively, (A,L(Z))1 is the supremum of standard deviations of Gaussians splittable off (A,L(Z))2; a structural lemma shows (A,L(Z))3 always, that (A,L(Z))4 is finite (a characteristic function decaying faster than every Gaussian would vanish off the origin), and that Gaussian-freeness is equivalent to (A,L(Z))5. The Gaussian splitting theorem then states that every real-valued random variable decomposes essentially uniquely (up to translation) as (A,L(Z))6 with (A,L(Z))7 Gaussian-free and (A,L(Z))8 independent. For infinitely divisible laws, (A,L(Z))9 equals the Gaussian coefficient X0 of the Lévy triplet.
Under the Gaussian-free hypothesis, the paper proves its strongest statement: with full column rank mixing matrices, the sources are identifiable up to permutation, scale and translation even in the presence of additive Gaussian noise with arbitrary, possibly non-diagonal and degenerate covariance. If both candidate source vectors are Gaussian-free, the residual Gaussian perturbation vanishes entirely and even the noise laws agree (X1), so the whole model X2 is identified. The paper emphasizes that part (1) without the two-sided hypothesis is genuinely one-sided — a Rademacher source plus unit Gaussian noise versus the noisy mixture itself satisfy all hypotheses of part (1) yet are not affinely related — and that the Gaussian-free assumption is a restriction on parametrization only, since every source can be split and its Gaussian part absorbed into X3. All standard super-Gaussian ICA source models (Laplace, Student, Cauchy) are Gaussian-free; a Rademacher-plus-Gaussian mixture is not, demonstrating the hypothesis is strictly stronger than non-Gaussianity.
Estimation: equivariant gradient descent in the complete noiseless model
For square invertible X4 and no noise, the classical identifiability corollary holds under the weakest hypothesis of at most one Gaussian source, with sources determined up to permutation and sign after centring and unit-variance normalization. Two Gaussian sources destroy identifiability completely — the ambiguity is then the whole orthogonal group.
On the algorithmic side, the paper sets up maximum likelihood, computes the gradient, and clarifies a point often stated incorrectly in the literature: the customary "preconditioner" X5 in the update is not an approximate inverse Hessian (the operator involved does not depend on the data, whereas the true Hessian does), but the exact gradient of the exact objective for the right-invariant Riemannian metric on X6 — Amari's natural gradient, Cardoso–Laheld's relative gradient. The update is equivariant: the dynamics lives on the global system matrix X7 and depends on X8 only through X9.
The stability analysis defines (1−∣t∣)+0 with (1−∣t∣)+1 at the fixed-point scale, and gives sufficient conditions for local asymptotic stability of the separating solution ((1−∣t∣)+2, (1−∣t∣)+3, (1−∣t∣)+4 for all pairs), with strict reversal implying instability for every step size. Three consequences deserve emphasis:
- Gaussian sources lie exactly on the boundary: for any model score whatsoever, Stein's identity gives (1−∣t∣)+5 for a Gaussian source, so two Gaussian sources force (1−∣t∣)+6 and instability of the neutral direction.
- With correctly specified scores, (1−∣t∣)+7 by Cauchy–Schwarz, with equality exactly at Gaussians; hence the stability conditions hold precisely when the identifiability theory declares the model identifiable. However, stable non-separating equilibria can exist even under perfect specification — the paper exhibits a concrete three-component Gaussian-mixture source law and a stationary point with all eigenvalues strictly negative that is not a generalized permutation matrix.
- For the cubic nonlinearity, (1−∣t∣)+8, so kurtosis sign is exactly the criterion. For (1−∣t∣)+9-based scores the criterion is a Stein discrepancy inequality, not the kurtosis sign: the paper gives counterexamples in both directions, including a four-atom law with kurtosis [−1,1]0 but [−1,1]1, for which two such super-Gaussian sources make the separating solution unstable under [−1,1]2.
The online algorithm is a Robbins–Monro scheme tracking the mean ODE; with constant step size it converges only to an [−1,1]3-neighbourhood of a stable equilibrium. Misspecified scores cost statistical efficiency (departure from the Cramér–Rao bound) but not local attraction, provided the stability inequalities hold.
LiNGAM and generalizations
Adding acyclicity and the unit-diagonal constraint removes the last ambiguities: the paper proves that LiNGAM's structural matrix [−1,1]4 and disturbances are identified exactly, via a cycle argument showing no non-trivial permutation preserves both unit diagonal and triangularizability. Each remaining assumption is mapped to the ambiguity it removes: the unit coefficient fixes scale, acyclicity kills permutation, independence of disturbances is causal sufficiency, and non-Gaussianity is what makes the whole argument available. A closing survey covers nonlinear ICA (unidentifiable in general by the Darmois construction; identifiable with auxiliary variables as in TCL, PCL, iVAE), multidimensional ICA, independent vector analysis, multi-view ICA, and additive-noise causal models.
Limitations and open questions
Several caveats are stated plainly. The stability theorem is local: the objective is not concave, and attracting non-separating equilibria exist even under correct specification, so global convergence from arbitrary initializations is not established. The borderline cases [−1,1]5 and [−1,1]6 are left undecided by the linearization. The regularity condition (bounded [−1,1]7) excludes the cubic nonlinearity, which requires a separate polynomial argument. The empirical objective requires stationarity/ergodicity for consistency — without it the criterion can be unbounded. On the identifiability side, the overcomplete case [−1,1]8 leaves [−1,1]9 without a left inverse, so sources cannot be recovered pointwise even when identifiable; and whether independent mechanism analysis alone delivers full identifiability of nonlinear mixtures remains open. Finally, the Gaussian-free framework identifies sources only up to translation, leaving open whether sharper canonical normalizations are achievable within it.
Conclusion
The paper consolidates the identifiability theory of linear ICA into a single self-contained development anchored on the Kagan–Linnik–Rao column dichotomy, extends it to a sharp Gaussian-free theory robust to arbitrary dependent Gaussian noise, and connects identifiability exactly to local stability of the equivariant gradient algorithm. Its corrections to folklore — that the relative-gradient preconditioner is a metric rather than a Hessian approximation, that kurtosis sign is not generally the right nonlinearity-selection rule, and that non-normality rather than non-zero kurtosis is the operative hypothesis — are of independent value to practitioners and theorists alike.
Paper to Video (Beta)
No one has generated a video about this paper yet.
Whiteboard
No one has generated a whiteboard explanation for this paper yet.
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Continue Learning
- How does the Kagan–Linnik–Rao theorem establish the column-by-column identifiability structure in linear ICA?
- Why is Gaussian-freeness a stronger and more useful condition than simple non-Gaussianity when additive noise is present?
- How do the stability conditions for equivariant gradient descent relate to Fisher information and source distributions?
- What are the practical consequences of separating kurtosis-based ICA assumptions from broader measures of non-Gaussianity?
- Find recent papers about identifiability and estimation in noisy independent component analysis.
Tweets
Sign up for free to view the 1 tweet with 75 likes about this paper.