Multiple and Inhomogeneous Component Analysis
- Multiple and Inhomogeneous Component Analysis (MICA) is a multi-view latent variable framework that isolates shared and view-specific components with varying distributions and transformations.
- It employs cumulant methods and noise diversity to achieve identifiability in both non-Gaussian and Gaussian settings.
- Algorithms such as RCA and ShICA-J leverage second- and higher-order statistics to accurately recover latent signals under complex mixing conditions.
Multiple and Inhomogeneous Component Analysis (MICA) denotes a class of multi-view latent-variable problems in which observed views are driven by multiple latent components that may be shared across only some views, may be absent from others, may have heterogeneous distributions, and may undergo view-specific transformations. In one formulation, MICA asks for the systematic separation and learning of latent components that drive multiple observed views, when components may differ across views, have heterogeneous distributions, and undergo view-specific transformations; Rich Component Analysis (RCA) provides a cumulant-based framework for isolating the cumulants of each latent component from multi-view observations without modeling the other components (Ge et al., 2015). In another formulation, MICA is the second-order identifiable regime of Shared Independent Component Analysis (ShICA), in which multiple views share the same latent independent sources and identifiability is driven by componentwise inhomogeneous Gaussian noise variances across views (Richard et al., 2021). Taken together, these formulations place MICA at the intersection of multi-view learning, latent-component identifiability, cumulant methods, multiset correlation methods, and independent component analysis.
1. Conceptual scope and problem formulation
MICA is “Multiple” because there are multiple views and multiple components, each present in different subsets of views; it is “Inhomogeneous” because each latent component may have a different distribution and different transformations across views, and some components may be absent from specific views (Ge et al., 2015). In the RCA formulation, the objective is not necessarily to reconstruct sample-level latent signals, but to separate and learn the distributions of latent components from overlapping multi-view mixtures. This directly addresses settings in which the observable data are complex mixtures and the nuisance components are not modeled parametrically (Ge et al., 2015).
A second, narrower formulation appears in ShICA, where MICA posits multiple views sharing the same latent independent sources , with view-specific mixing matrices and per-component, per-view Gaussian noise variances . The hallmark of this regime is that the inhomogeneous variance patterns across views differ componentwise, enabling second-order identification of components even when the sources are Gaussian (Richard et al., 2021).
These two formulations emphasize different sources of identifiability. RCA exploits additivity of cumulants under independence and higher-order non-Gaussian structure to isolate components in very general multi-view mixtures (Ge et al., 2015). ShICA exploits “noise diversity” across views, so that second-order structure alone can identify shared components in the Gaussian or mixed Gaussian/non-Gaussian case under its model assumptions (Richard et al., 2021). This suggests that MICA is best understood as a family of identifiability problems rather than as a single fixed model.
2. Generative models and notation
In the RCA formulation, let index views and let each observed random vector be generated by a subset of latent components , where each latent is associated with a subset of views. Components are mutually independent, and the model uses view-specific linear maps , with the zero map when component is absent from view 0. The generative model is
1
In the core RCA analysis, the noise term 2 is either absorbed into a view-specific component or handled via robustness bounds; most theory is stated for the noiseless latent linear mixture (Ge et al., 2015).
An instructive special case is the two-view contrastive model
3
where 4 is unique to 5, 6 unique to 7, and 8 is shared across the two views with unknown linear transform 9. All 0 are independent and may be complex, high-dimensional, and non-Gaussian. Means are assumed 1 or known, which is without loss because cumulant extraction then determines all higher cumulants (Ge et al., 2015).
In ShICA, the formal model is
2
where 3 is invertible, 4 has independent components with unit variance, and 5 is view-specific additive noise independent of 6. The noise is added in the component domain: 7 with 8. Rewriting this as 9 with sensor-domain noise 0 is possible algebraically, but the ShICA identifiability theory relies on the component-domain diagonal-noise structure (Richard et al., 2021).
The notational correspondence between the RCA and MICA descriptions is explicit in the supplied material: RCA uses 1 for views and 2 for components, whereas the MICA description uses 3 for views and the same 4 for latent components; RCA’s incidence sets 5 correspond to the view-component incidence structure in MICA, and RCA’s linear maps 6 correspond to linearized view-specific transformations 7 (Ge et al., 2015).
3. Cumulant-based component isolation in Rich Component Analysis
RCA is built on cumulants and cross-cumulants. For a scalar random variable,
8
is the cumulant generating function, and the 9-th cumulant is 0 times the coefficient of 1 in 2. For 3,
4
and the 5-th cumulant tensor 6 has entries
7
The multivariate cross-cumulant is
8
with the sum over all set partitions 9 of 0 (Ge et al., 2015).
The key property is additivity under independence: 1 whenever 2 are independent of 3. RCA uses this property to separate latent components additively, without specifying nuisance distributions (Ge et al., 2015).
For centered vector data,
4
5
and
6
For Gaussian variables, 7 for all 8; RCA exploits this to ignore Gaussian nuisance components via higher-order cumulants (Ge et al., 2015).
In the two-view contrastive setting, RCA derives explicit identities for recovering the shared transform and the component cumulants. Let 9 be the standard unfolding along the last mode. Then
0
so
1
This works because only 2 appears in both 3 and 4; full rank of 5 is required, and this condition is not satisfied by Gaussian 6 since 7 (Ge et al., 2015).
Once 8 is known, RCA isolates the shared component through
9
and therefore
0
An equivalent sum-of-views identity avoids mixed cumulant computation: 1 If 2 has higher dimension than 3, 4 is replaced by the Moore–Penrose pseudoinverse 5 (Ge et al., 2015).
The general multi-view formulation extends these identities. The model is
6
where 7 appears in views indexed by 8, and identifiability requires all 9 to be distinct. RCA introduces the notion of an 0-distinguishable set system: a family 1 is 2-distinguishable if every 3 has a distinguishing subset 4 of size at most 5 such that for any 6, either 7 or 8. This condition guarantees that appropriate cross-cumulants isolate 9, up to supersets that are handled recursively (Ge et al., 2015).
4. Second-order identifiability and the ShICA interpretation of MICA
ShICA gives a second-order formulation of MICA in which identifiability is driven by inhomogeneous per-component Gaussian noise variances across views (Richard et al., 2021). The source prior factorizes as
0
and some components may be Gaussian while others are non-Gaussian. The main identifiability theorem assumes 1 views and requires a noise-diversity condition for Gaussian components: for all 2 with 3, the sequences 4 and 5 are different (Richard et al., 2021).
The cross-covariances are
6
In particular, for 7,
8
since the component-domain noises are independent across views (Richard et al., 2021).
Theorem 1 in ShICA states that if 9 and the noise-diversity condition holds, then two parameter sets generating the same joint distribution differ only by the usual ICA indeterminacies: there exists a sign-permutation matrix 00 such that for all 01,
02
Thus the model is globally identifiable up to permutation and sign (Richard et al., 2021).
When all components are Gaussian, second-order statistics suffice. Let
03
form the block matrices
04
and consider the generalized eigenproblem
05
Under the assumption that the top 06 eigenvalues 07 are distinct, Theorem 2 states that there exists a permutation 08 and per-view diagonal scale matrices 09 such that
10
where 11 is obtained from the top generalized eigenvectors (Richard et al., 2021).
The eigenvalues themselves are characterized componentwise: for each component 12, 13 is the largest solution of
14
and the remaining eigenvalues are 15. Distinct eigenvalues therefore require distinct variance sequences 16, linking the generalized-eigenvalue criterion directly to inhomogeneous noise patterns across views (Richard et al., 2021).
A notable subtlety is that global identifiability and direct recovery by Multiset CCA are not equivalent. ShICA gives a counterexample in which two components have the same multiset of variances across views, so the model remains globally identifiable because the noise-diversity condition still holds, but 17, preventing separation by eigenvectors alone (Richard et al., 2021). This distinction is central to the ShICA interpretation of MICA: the model may be identifiable even when a particular second-order algorithm is not sufficient by itself.
5. Algorithms and estimation procedures
RCA provides two named algorithms for the general multi-view setting: FindLinear and ComputeCumulant. FindLinear iterates over maximal sets 18, uses cumulants of order 19 on distinguishing sets 20, subtracts already recovered superset contributions, and estimates
21
with
22
ComputeCumulant then uses the recovered 23 to obtain 24 for 25 by inverse transforms and subtraction of superset contributions (Ge et al., 2015).
For the two-view model, the algorithmic workflow is simpler: estimate 26 from fourth-order cross-cumulants; extract 27, 28, and 29; and then learn the target component using method-of-moments or a cumulant-corrected stochastic-gradient procedure (Ge et al., 2015). The RCA meta-algorithm uses extracted cumulants or moments to form unbiased or low-bias estimates of the gradient of the target log-likelihood, even though only perturbed multi-view samples are observed. If 30 is 31-strongly convex and 32-smooth and if the gradient estimator 33 satisfies
34
then gradient descent with step size 35 yields
36
This guarantee makes cumulant-corrected SGD a bridge between component isolation and standard optimization theory (Ge et al., 2015).
The ShICA-J algorithm is the principal second-order estimator in the ShICA formulation of MICA. It begins with Multiset CCA: estimate the block covariances 37, assemble 38 and 39, and solve the generalized eigenproblem 40. The resulting generalized eigenvectors define matrices 41 that recover the unmixing up to permutation and diagonal scaling when the eigenvalues are distinct (Richard et al., 2021).
Because sampling noise perturbs the block covariances and generalized eigenvectors can be unstable when eigen-gaps are small, ShICA-J adds a joint diagonalization step. It computes
42
and finds a rotation 43 by minimizing
44
The final subspace-aligned unmixing is 45 (Richard et al., 2021). Per-view scaling is then recovered from
46
through the optimization
47
using the fixed-point updates stated in the supplied material. The final ShICA-J unmixing matrices are
48
ShICA-J is described as fast and robust to sampling noise under MICA conditions (Richard et al., 2021).
ShICA-ML extends this second-order regime by introducing a maximum-likelihood estimator that leverages non-Gaussianity. For non-Gaussian components, it uses a super-Gaussian mixture-of-Gaussians prior
49
optimizes the likelihood with a generalized EM procedure, uses closed-form E-steps for posterior statistics of the components, and updates 50 by a quasi-Newton step with the relative gradient and approximate Hessian given in the supplied material (Richard et al., 2021). In the Gaussian case, ShICA also provides an MMSE estimator for the shared components: 51 with posterior covariance
52
This explicit fusion rule is presented as a principled method for shared-components estimation (Richard et al., 2021).
6. Assumptions, comparisons, empirical behavior, and limitations
The central assumptions differ between the two MICA formulations. RCA assumes mutual independence of the latent components, existence of moments and cumulants up to the required order, invertibility of the nonzero transformation matrices, and an 53-distinguishable set system for the multi-view incidence structure (Ge et al., 2015). Non-Gaussianity is important for identifiability through higher-order cumulants, because Gaussian components satisfy 54 for all 55. If a component is Gaussian, it disappears from higher-order cumulants, and RCA exploits this by focusing on orders where the target is nonzero (Ge et al., 2015).
ShICA assumes component-domain additive Gaussian noise with diagonal covariance, independent sources of unit variance, and 56 for the main identifiability theorem. Its second-order identifiability requires unique variance fingerprints across views for Gaussian components. Identifiability fails when all components are Gaussian and variance fingerprints across views are identical, since rotations within the Gaussian subspace are then not broken. For 57, there is single-view rotational indeterminacy; for 58, stronger conditions are required and recovery is only up to scaling (Richard et al., 2021).
The comparison with related methods is explicit in the supplied material. Relative to ICA, RCA/MICA allows each component 59 to be high-dimensional and arbitrarily complex, works across multiple views and subset-sharing patterns, and aims to learn component distributions through cumulants rather than to recover each sample-level signal (Ge et al., 2015). Relative to CCA, RCA does not assume Gaussianity and can extract cumulants of non-Gaussian components even when they overlap with shared subspaces; the supplied material states that RCA empirically outperforms naive CCA-based projection in contrastive tasks because CCA only captures second-order shared covariance structure (Ge et al., 2015). In ShICA, Multiset CCA can recover the correct unmixing matrices in some cases, but even a small amount of sampling noise makes Multiset CCA fail; the ShICA-J correction addresses this by combining MCCA with joint diagonalization (Richard et al., 2021).
The empirical record reported in the supplied material is also bifurcated. For RCA, synthetic tasks include PCA, linear regression, mixture of Gaussians, logistic regression, and Ising model contrastive learning; RCA consistently outperformed naive and CCA baselines, and with approximately 60 samples approached the “true samples” gold standard in mean squared error across tasks (Ge et al., 2015). In a real-data biomarker example with DNA methylation markers and simulated lab-induced perturbations, RCA reduced the MSE of logistic-regression coefficient estimates to approximately 61, compared with approximately 62 for naive estimation, approximately 63 with CCA, and approximately 64 when using control markers as covariates (Ge et al., 2015).
For ShICA, the reported behavior depends on whether the data satisfy the MICA regime or require higher-order non-Gaussian modeling. In Gaussian-only settings with variance diversity, ShICA-J and ShICA-ML separate well, whereas MCCA alone needs many samples and is sensitive to noise; in non-Gaussian-only settings without variance diversity, ShICA-ML and CanICA separate, while ShICA-J and MCCA cannot because second-order information is insufficient; in mixed Gaussian/non-Gaussian settings, ShICA-ML performs best, while second-order methods do not fully separate the components (Richard et al., 2021). On fMRI reconstruction of left-out subjects, ShICA-ML is reported to give the highest 65 across several datasets, with ShICA-J competitive and faster. On MEG data from Cam-CAN, ShICA-ML yields significantly lower across-trial variability of the recovered shared components, while ShICA-J remains competitive and much faster (Richard et al., 2021).
The principal limitations are likewise explicit. RCA’s higher-order cumulant estimation has exponential-in-order sample complexity; full-rank tensor-unfolding conditions are required; independence is essential because cumulant additivity breaks under correlated components; and the cumulant order must be chosen carefully, since symmetric distributions may have vanishing third cumulant (Ge et al., 2015). ShICA does not reduce dimensionality by itself, so external per-view PCA is needed in very high-dimensional settings; its guarantees do not apply directly under non-Gaussian noise, dependent sources, or sensor-domain noise without the diagonal component-domain structure (Richard et al., 2021). A plausible implication is that “MICA” names a broad methodological objective, but the practical form of identifiability depends sharply on whether one exploits higher-order cumulants in heterogeneous multi-component mixtures or second-order variance diversity in shared-source models.