Papers
Topics
Authors
Recent
Search
2000 character limit reached

Multiple and Inhomogeneous Component Analysis

Updated 10 July 2026
  • Multiple and Inhomogeneous Component Analysis (MICA) is a multi-view latent variable framework that isolates shared and view-specific components with varying distributions and transformations.
  • It employs cumulant methods and noise diversity to achieve identifiability in both non-Gaussian and Gaussian settings.
  • Algorithms such as RCA and ShICA-J leverage second- and higher-order statistics to accurately recover latent signals under complex mixing conditions.

Multiple and Inhomogeneous Component Analysis (MICA) denotes a class of multi-view latent-variable problems in which observed views are driven by multiple latent components that may be shared across only some views, may be absent from others, may have heterogeneous distributions, and may undergo view-specific transformations. In one formulation, MICA asks for the systematic separation and learning of latent components that drive multiple observed views, when components may differ across views, have heterogeneous distributions, and undergo view-specific transformations; Rich Component Analysis (RCA) provides a cumulant-based framework for isolating the cumulants of each latent component from multi-view observations without modeling the other components (Ge et al., 2015). In another formulation, MICA is the second-order identifiable regime of Shared Independent Component Analysis (ShICA), in which multiple views share the same latent independent sources and identifiability is driven by componentwise inhomogeneous Gaussian noise variances across views (Richard et al., 2021). Taken together, these formulations place MICA at the intersection of multi-view learning, latent-component identifiability, cumulant methods, multiset correlation methods, and independent component analysis.

1. Conceptual scope and problem formulation

MICA is “Multiple” because there are multiple views and multiple components, each present in different subsets of views; it is “Inhomogeneous” because each latent component may have a different distribution and different transformations across views, and some components may be absent from specific views (Ge et al., 2015). In the RCA formulation, the objective is not necessarily to reconstruct sample-level latent signals, but to separate and learn the distributions of latent components from overlapping multi-view mixtures. This directly addresses settings in which the observable data are complex mixtures and the nuisance components are not modeled parametrically (Ge et al., 2015).

A second, narrower formulation appears in ShICA, where MICA posits multiple views sharing the same latent independent sources ss, with view-specific mixing matrices A(v)A^{(v)} and per-component, per-view Gaussian noise variances Σ(v)=diag(σ1,v2,,σp,v2)\Sigma^{(v)} = \mathrm{diag}(\sigma_{1,v}^2,\ldots,\sigma_{p,v}^2). The hallmark of this regime is that the inhomogeneous variance patterns across views differ componentwise, enabling second-order identification of components even when the sources are Gaussian (Richard et al., 2021).

These two formulations emphasize different sources of identifiability. RCA exploits additivity of cumulants under independence and higher-order non-Gaussian structure to isolate components in very general multi-view mixtures (Ge et al., 2015). ShICA exploits “noise diversity” across views, so that second-order structure alone can identify shared components in the Gaussian or mixed Gaussian/non-Gaussian case under its model assumptions (Richard et al., 2021). This suggests that MICA is best understood as a family of identifiability problems rather than as a single fixed model.

2. Generative models and notation

In the RCA formulation, let v{1,,V}v \in \{1,\ldots,V\} index views and let each observed random vector X(v)RdX^{(v)} \in \mathbb{R}^d be generated by a subset of latent components {Sj}j=1p\{S_j\}_{j=1}^p, where each latent SjS_j is associated with a subset Qj{1,,V}Q_j \subseteq \{1,\ldots,V\} of views. Components are mutually independent, and the model uses view-specific linear maps A(v,j)Rd×dA^{(v,j)} \in \mathbb{R}^{d \times d}, with the zero map when component jj is absent from view A(v)A^{(v)}0. The generative model is

A(v)A^{(v)}1

In the core RCA analysis, the noise term A(v)A^{(v)}2 is either absorbed into a view-specific component or handled via robustness bounds; most theory is stated for the noiseless latent linear mixture (Ge et al., 2015).

An instructive special case is the two-view contrastive model

A(v)A^{(v)}3

where A(v)A^{(v)}4 is unique to A(v)A^{(v)}5, A(v)A^{(v)}6 unique to A(v)A^{(v)}7, and A(v)A^{(v)}8 is shared across the two views with unknown linear transform A(v)A^{(v)}9. All Σ(v)=diag(σ1,v2,,σp,v2)\Sigma^{(v)} = \mathrm{diag}(\sigma_{1,v}^2,\ldots,\sigma_{p,v}^2)0 are independent and may be complex, high-dimensional, and non-Gaussian. Means are assumed Σ(v)=diag(σ1,v2,,σp,v2)\Sigma^{(v)} = \mathrm{diag}(\sigma_{1,v}^2,\ldots,\sigma_{p,v}^2)1 or known, which is without loss because cumulant extraction then determines all higher cumulants (Ge et al., 2015).

In ShICA, the formal model is

Σ(v)=diag(σ1,v2,,σp,v2)\Sigma^{(v)} = \mathrm{diag}(\sigma_{1,v}^2,\ldots,\sigma_{p,v}^2)2

where Σ(v)=diag(σ1,v2,,σp,v2)\Sigma^{(v)} = \mathrm{diag}(\sigma_{1,v}^2,\ldots,\sigma_{p,v}^2)3 is invertible, Σ(v)=diag(σ1,v2,,σp,v2)\Sigma^{(v)} = \mathrm{diag}(\sigma_{1,v}^2,\ldots,\sigma_{p,v}^2)4 has independent components with unit variance, and Σ(v)=diag(σ1,v2,,σp,v2)\Sigma^{(v)} = \mathrm{diag}(\sigma_{1,v}^2,\ldots,\sigma_{p,v}^2)5 is view-specific additive noise independent of Σ(v)=diag(σ1,v2,,σp,v2)\Sigma^{(v)} = \mathrm{diag}(\sigma_{1,v}^2,\ldots,\sigma_{p,v}^2)6. The noise is added in the component domain: Σ(v)=diag(σ1,v2,,σp,v2)\Sigma^{(v)} = \mathrm{diag}(\sigma_{1,v}^2,\ldots,\sigma_{p,v}^2)7 with Σ(v)=diag(σ1,v2,,σp,v2)\Sigma^{(v)} = \mathrm{diag}(\sigma_{1,v}^2,\ldots,\sigma_{p,v}^2)8. Rewriting this as Σ(v)=diag(σ1,v2,,σp,v2)\Sigma^{(v)} = \mathrm{diag}(\sigma_{1,v}^2,\ldots,\sigma_{p,v}^2)9 with sensor-domain noise v{1,,V}v \in \{1,\ldots,V\}0 is possible algebraically, but the ShICA identifiability theory relies on the component-domain diagonal-noise structure (Richard et al., 2021).

The notational correspondence between the RCA and MICA descriptions is explicit in the supplied material: RCA uses v{1,,V}v \in \{1,\ldots,V\}1 for views and v{1,,V}v \in \{1,\ldots,V\}2 for components, whereas the MICA description uses v{1,,V}v \in \{1,\ldots,V\}3 for views and the same v{1,,V}v \in \{1,\ldots,V\}4 for latent components; RCA’s incidence sets v{1,,V}v \in \{1,\ldots,V\}5 correspond to the view-component incidence structure in MICA, and RCA’s linear maps v{1,,V}v \in \{1,\ldots,V\}6 correspond to linearized view-specific transformations v{1,,V}v \in \{1,\ldots,V\}7 (Ge et al., 2015).

3. Cumulant-based component isolation in Rich Component Analysis

RCA is built on cumulants and cross-cumulants. For a scalar random variable,

v{1,,V}v \in \{1,\ldots,V\}8

is the cumulant generating function, and the v{1,,V}v \in \{1,\ldots,V\}9-th cumulant is X(v)RdX^{(v)} \in \mathbb{R}^d0 times the coefficient of X(v)RdX^{(v)} \in \mathbb{R}^d1 in X(v)RdX^{(v)} \in \mathbb{R}^d2. For X(v)RdX^{(v)} \in \mathbb{R}^d3,

X(v)RdX^{(v)} \in \mathbb{R}^d4

and the X(v)RdX^{(v)} \in \mathbb{R}^d5-th cumulant tensor X(v)RdX^{(v)} \in \mathbb{R}^d6 has entries

X(v)RdX^{(v)} \in \mathbb{R}^d7

The multivariate cross-cumulant is

X(v)RdX^{(v)} \in \mathbb{R}^d8

with the sum over all set partitions X(v)RdX^{(v)} \in \mathbb{R}^d9 of {Sj}j=1p\{S_j\}_{j=1}^p0 (Ge et al., 2015).

The key property is additivity under independence: {Sj}j=1p\{S_j\}_{j=1}^p1 whenever {Sj}j=1p\{S_j\}_{j=1}^p2 are independent of {Sj}j=1p\{S_j\}_{j=1}^p3. RCA uses this property to separate latent components additively, without specifying nuisance distributions (Ge et al., 2015).

For centered vector data,

{Sj}j=1p\{S_j\}_{j=1}^p4

{Sj}j=1p\{S_j\}_{j=1}^p5

and

{Sj}j=1p\{S_j\}_{j=1}^p6

For Gaussian variables, {Sj}j=1p\{S_j\}_{j=1}^p7 for all {Sj}j=1p\{S_j\}_{j=1}^p8; RCA exploits this to ignore Gaussian nuisance components via higher-order cumulants (Ge et al., 2015).

In the two-view contrastive setting, RCA derives explicit identities for recovering the shared transform and the component cumulants. Let {Sj}j=1p\{S_j\}_{j=1}^p9 be the standard unfolding along the last mode. Then

SjS_j0

so

SjS_j1

This works because only SjS_j2 appears in both SjS_j3 and SjS_j4; full rank of SjS_j5 is required, and this condition is not satisfied by Gaussian SjS_j6 since SjS_j7 (Ge et al., 2015).

Once SjS_j8 is known, RCA isolates the shared component through

SjS_j9

and therefore

Qj{1,,V}Q_j \subseteq \{1,\ldots,V\}0

An equivalent sum-of-views identity avoids mixed cumulant computation: Qj{1,,V}Q_j \subseteq \{1,\ldots,V\}1 If Qj{1,,V}Q_j \subseteq \{1,\ldots,V\}2 has higher dimension than Qj{1,,V}Q_j \subseteq \{1,\ldots,V\}3, Qj{1,,V}Q_j \subseteq \{1,\ldots,V\}4 is replaced by the Moore–Penrose pseudoinverse Qj{1,,V}Q_j \subseteq \{1,\ldots,V\}5 (Ge et al., 2015).

The general multi-view formulation extends these identities. The model is

Qj{1,,V}Q_j \subseteq \{1,\ldots,V\}6

where Qj{1,,V}Q_j \subseteq \{1,\ldots,V\}7 appears in views indexed by Qj{1,,V}Q_j \subseteq \{1,\ldots,V\}8, and identifiability requires all Qj{1,,V}Q_j \subseteq \{1,\ldots,V\}9 to be distinct. RCA introduces the notion of an A(v,j)Rd×dA^{(v,j)} \in \mathbb{R}^{d \times d}0-distinguishable set system: a family A(v,j)Rd×dA^{(v,j)} \in \mathbb{R}^{d \times d}1 is A(v,j)Rd×dA^{(v,j)} \in \mathbb{R}^{d \times d}2-distinguishable if every A(v,j)Rd×dA^{(v,j)} \in \mathbb{R}^{d \times d}3 has a distinguishing subset A(v,j)Rd×dA^{(v,j)} \in \mathbb{R}^{d \times d}4 of size at most A(v,j)Rd×dA^{(v,j)} \in \mathbb{R}^{d \times d}5 such that for any A(v,j)Rd×dA^{(v,j)} \in \mathbb{R}^{d \times d}6, either A(v,j)Rd×dA^{(v,j)} \in \mathbb{R}^{d \times d}7 or A(v,j)Rd×dA^{(v,j)} \in \mathbb{R}^{d \times d}8. This condition guarantees that appropriate cross-cumulants isolate A(v,j)Rd×dA^{(v,j)} \in \mathbb{R}^{d \times d}9, up to supersets that are handled recursively (Ge et al., 2015).

4. Second-order identifiability and the ShICA interpretation of MICA

ShICA gives a second-order formulation of MICA in which identifiability is driven by inhomogeneous per-component Gaussian noise variances across views (Richard et al., 2021). The source prior factorizes as

jj0

and some components may be Gaussian while others are non-Gaussian. The main identifiability theorem assumes jj1 views and requires a noise-diversity condition for Gaussian components: for all jj2 with jj3, the sequences jj4 and jj5 are different (Richard et al., 2021).

The cross-covariances are

jj6

In particular, for jj7,

jj8

since the component-domain noises are independent across views (Richard et al., 2021).

Theorem 1 in ShICA states that if jj9 and the noise-diversity condition holds, then two parameter sets generating the same joint distribution differ only by the usual ICA indeterminacies: there exists a sign-permutation matrix A(v)A^{(v)}00 such that for all A(v)A^{(v)}01,

A(v)A^{(v)}02

Thus the model is globally identifiable up to permutation and sign (Richard et al., 2021).

When all components are Gaussian, second-order statistics suffice. Let

A(v)A^{(v)}03

form the block matrices

A(v)A^{(v)}04

and consider the generalized eigenproblem

A(v)A^{(v)}05

Under the assumption that the top A(v)A^{(v)}06 eigenvalues A(v)A^{(v)}07 are distinct, Theorem 2 states that there exists a permutation A(v)A^{(v)}08 and per-view diagonal scale matrices A(v)A^{(v)}09 such that

A(v)A^{(v)}10

where A(v)A^{(v)}11 is obtained from the top generalized eigenvectors (Richard et al., 2021).

The eigenvalues themselves are characterized componentwise: for each component A(v)A^{(v)}12, A(v)A^{(v)}13 is the largest solution of

A(v)A^{(v)}14

and the remaining eigenvalues are A(v)A^{(v)}15. Distinct eigenvalues therefore require distinct variance sequences A(v)A^{(v)}16, linking the generalized-eigenvalue criterion directly to inhomogeneous noise patterns across views (Richard et al., 2021).

A notable subtlety is that global identifiability and direct recovery by Multiset CCA are not equivalent. ShICA gives a counterexample in which two components have the same multiset of variances across views, so the model remains globally identifiable because the noise-diversity condition still holds, but A(v)A^{(v)}17, preventing separation by eigenvectors alone (Richard et al., 2021). This distinction is central to the ShICA interpretation of MICA: the model may be identifiable even when a particular second-order algorithm is not sufficient by itself.

5. Algorithms and estimation procedures

RCA provides two named algorithms for the general multi-view setting: FindLinear and ComputeCumulant. FindLinear iterates over maximal sets A(v)A^{(v)}18, uses cumulants of order A(v)A^{(v)}19 on distinguishing sets A(v)A^{(v)}20, subtracts already recovered superset contributions, and estimates

A(v)A^{(v)}21

with

A(v)A^{(v)}22

ComputeCumulant then uses the recovered A(v)A^{(v)}23 to obtain A(v)A^{(v)}24 for A(v)A^{(v)}25 by inverse transforms and subtraction of superset contributions (Ge et al., 2015).

For the two-view model, the algorithmic workflow is simpler: estimate A(v)A^{(v)}26 from fourth-order cross-cumulants; extract A(v)A^{(v)}27, A(v)A^{(v)}28, and A(v)A^{(v)}29; and then learn the target component using method-of-moments or a cumulant-corrected stochastic-gradient procedure (Ge et al., 2015). The RCA meta-algorithm uses extracted cumulants or moments to form unbiased or low-bias estimates of the gradient of the target log-likelihood, even though only perturbed multi-view samples are observed. If A(v)A^{(v)}30 is A(v)A^{(v)}31-strongly convex and A(v)A^{(v)}32-smooth and if the gradient estimator A(v)A^{(v)}33 satisfies

A(v)A^{(v)}34

then gradient descent with step size A(v)A^{(v)}35 yields

A(v)A^{(v)}36

This guarantee makes cumulant-corrected SGD a bridge between component isolation and standard optimization theory (Ge et al., 2015).

The ShICA-J algorithm is the principal second-order estimator in the ShICA formulation of MICA. It begins with Multiset CCA: estimate the block covariances A(v)A^{(v)}37, assemble A(v)A^{(v)}38 and A(v)A^{(v)}39, and solve the generalized eigenproblem A(v)A^{(v)}40. The resulting generalized eigenvectors define matrices A(v)A^{(v)}41 that recover the unmixing up to permutation and diagonal scaling when the eigenvalues are distinct (Richard et al., 2021).

Because sampling noise perturbs the block covariances and generalized eigenvectors can be unstable when eigen-gaps are small, ShICA-J adds a joint diagonalization step. It computes

A(v)A^{(v)}42

and finds a rotation A(v)A^{(v)}43 by minimizing

A(v)A^{(v)}44

The final subspace-aligned unmixing is A(v)A^{(v)}45 (Richard et al., 2021). Per-view scaling is then recovered from

A(v)A^{(v)}46

through the optimization

A(v)A^{(v)}47

using the fixed-point updates stated in the supplied material. The final ShICA-J unmixing matrices are

A(v)A^{(v)}48

ShICA-J is described as fast and robust to sampling noise under MICA conditions (Richard et al., 2021).

ShICA-ML extends this second-order regime by introducing a maximum-likelihood estimator that leverages non-Gaussianity. For non-Gaussian components, it uses a super-Gaussian mixture-of-Gaussians prior

A(v)A^{(v)}49

optimizes the likelihood with a generalized EM procedure, uses closed-form E-steps for posterior statistics of the components, and updates A(v)A^{(v)}50 by a quasi-Newton step with the relative gradient and approximate Hessian given in the supplied material (Richard et al., 2021). In the Gaussian case, ShICA also provides an MMSE estimator for the shared components: A(v)A^{(v)}51 with posterior covariance

A(v)A^{(v)}52

This explicit fusion rule is presented as a principled method for shared-components estimation (Richard et al., 2021).

6. Assumptions, comparisons, empirical behavior, and limitations

The central assumptions differ between the two MICA formulations. RCA assumes mutual independence of the latent components, existence of moments and cumulants up to the required order, invertibility of the nonzero transformation matrices, and an A(v)A^{(v)}53-distinguishable set system for the multi-view incidence structure (Ge et al., 2015). Non-Gaussianity is important for identifiability through higher-order cumulants, because Gaussian components satisfy A(v)A^{(v)}54 for all A(v)A^{(v)}55. If a component is Gaussian, it disappears from higher-order cumulants, and RCA exploits this by focusing on orders where the target is nonzero (Ge et al., 2015).

ShICA assumes component-domain additive Gaussian noise with diagonal covariance, independent sources of unit variance, and A(v)A^{(v)}56 for the main identifiability theorem. Its second-order identifiability requires unique variance fingerprints across views for Gaussian components. Identifiability fails when all components are Gaussian and variance fingerprints across views are identical, since rotations within the Gaussian subspace are then not broken. For A(v)A^{(v)}57, there is single-view rotational indeterminacy; for A(v)A^{(v)}58, stronger conditions are required and recovery is only up to scaling (Richard et al., 2021).

The comparison with related methods is explicit in the supplied material. Relative to ICA, RCA/MICA allows each component A(v)A^{(v)}59 to be high-dimensional and arbitrarily complex, works across multiple views and subset-sharing patterns, and aims to learn component distributions through cumulants rather than to recover each sample-level signal (Ge et al., 2015). Relative to CCA, RCA does not assume Gaussianity and can extract cumulants of non-Gaussian components even when they overlap with shared subspaces; the supplied material states that RCA empirically outperforms naive CCA-based projection in contrastive tasks because CCA only captures second-order shared covariance structure (Ge et al., 2015). In ShICA, Multiset CCA can recover the correct unmixing matrices in some cases, but even a small amount of sampling noise makes Multiset CCA fail; the ShICA-J correction addresses this by combining MCCA with joint diagonalization (Richard et al., 2021).

The empirical record reported in the supplied material is also bifurcated. For RCA, synthetic tasks include PCA, linear regression, mixture of Gaussians, logistic regression, and Ising model contrastive learning; RCA consistently outperformed naive and CCA baselines, and with approximately A(v)A^{(v)}60 samples approached the “true samples” gold standard in mean squared error across tasks (Ge et al., 2015). In a real-data biomarker example with DNA methylation markers and simulated lab-induced perturbations, RCA reduced the MSE of logistic-regression coefficient estimates to approximately A(v)A^{(v)}61, compared with approximately A(v)A^{(v)}62 for naive estimation, approximately A(v)A^{(v)}63 with CCA, and approximately A(v)A^{(v)}64 when using control markers as covariates (Ge et al., 2015).

For ShICA, the reported behavior depends on whether the data satisfy the MICA regime or require higher-order non-Gaussian modeling. In Gaussian-only settings with variance diversity, ShICA-J and ShICA-ML separate well, whereas MCCA alone needs many samples and is sensitive to noise; in non-Gaussian-only settings without variance diversity, ShICA-ML and CanICA separate, while ShICA-J and MCCA cannot because second-order information is insufficient; in mixed Gaussian/non-Gaussian settings, ShICA-ML performs best, while second-order methods do not fully separate the components (Richard et al., 2021). On fMRI reconstruction of left-out subjects, ShICA-ML is reported to give the highest A(v)A^{(v)}65 across several datasets, with ShICA-J competitive and faster. On MEG data from Cam-CAN, ShICA-ML yields significantly lower across-trial variability of the recovered shared components, while ShICA-J remains competitive and much faster (Richard et al., 2021).

The principal limitations are likewise explicit. RCA’s higher-order cumulant estimation has exponential-in-order sample complexity; full-rank tensor-unfolding conditions are required; independence is essential because cumulant additivity breaks under correlated components; and the cumulant order must be chosen carefully, since symmetric distributions may have vanishing third cumulant (Ge et al., 2015). ShICA does not reduce dimensionality by itself, so external per-view PCA is needed in very high-dimensional settings; its guarantees do not apply directly under non-Gaussian noise, dependent sources, or sensor-domain noise without the diagonal component-domain structure (Richard et al., 2021). A plausible implication is that “MICA” names a broad methodological objective, but the practical form of identifiability depends sharply on whether one exploits higher-order cumulants in heterogeneous multi-component mixtures or second-order variance diversity in shared-source models.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (2)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Multiple and Inhomogeneous Component Analysis (MICA).