Papers
Topics
Authors
Recent
Search
2000 character limit reached

MFMDScen in the MMGFM Framework

Updated 5 July 2026
  • MFMDScen is an informal shorthand linked to MMGFM, highlighting simulation scenarios rather than defining a distinct method.
  • The MMGFM model integrates multiple studies and modalities by separating shared and study-specific latent factors with covariate adjustments.
  • Simulation scenarios demonstrate MMGFM’s robustness and superior performance in handling complex, mixed-type data compared to alternative methods.

Within the context of (Liu et al., 14 Jul 2025), MFMDScen is not a defined acronym. The source paper explicitly introduces MMGFM, a high-dimensional multi-study, multi-modality, covariate-augmented generalized factor model, and uses “scenario” only for the simulation section’s Scenarios 1–4. Accordingly, the technically grounded interpretation of MFMDScen is as an apparent mislabeling or informal shorthand connected to the MMGFM framework rather than a distinct method. MMGFM is designed for settings with multiple studies, multiple modalities, mixed variable types, and additional covariates, and it combines study-shared, study-specific, and modality-related latent structure in a single generalized factor-analysis model (Liu et al., 14 Jul 2025).

1. Terminological status and scope

The cited work does not define the term MFMDScen. Instead, it defines MMGFM and frames its contribution as addressing a gap left by methods that focus predominantly either on multi-study integration or on multi-modality integration, but not on data with diverse modalities measured across multiple studies simultaneously (Liu et al., 14 Jul 2025).

This distinction matters for interpretation. If MFMDScen is encountered in connection with this paper, the evidence supports two constrained readings. First, it may refer indirectly to MMGFM itself. Second, it may refer to the paper’s simulation Scenarios 1–4, since “scenario” language appears there but not as a named acronym. A plausible implication is that MFMDScen should not be treated as an established methodological term unless further source material defines it explicitly.

2. Data structure and model specification

MMGFM is formulated for data from multiple studies s=1,,Ss=1,\dots,S, each containing multiple modalities m=1,,Mm=1,\dots,M, with potentially different variable types across modalities. For study ss and modality mm, the observed matrix is

Xsm=(xsimj)ins,  jpmRns×pm,X_{sm}=(x_{simj})_{i\le n_s,\; j\le p_m}\in\mathbb{R}^{n_s\times p_m},

with covariates zsiRdz_{si}\in\mathbb{R}^d. The hierarchical generalized factor-analysis form is

xsimjysimjc.i.d.EF(gm(ysimj)),ysimj=τsim+zsiβmj+λmjfi+γsmjgsi+vsim+εsimj.x_{simj}\mid y_{simj} \stackrel{c.i.d.}{\sim} EF(g_m(y_{simj})),\qquad y_{simj} = \tau_{sim} + z_{si}^\top \beta_{mj} + \lambda_{mj}^\top f_i + \gamma_{smj}^\top g_{si} + v_{sim} + \varepsilon_{simj}.

The latent decomposition is organized as follows:

  • fif_i: study-shared latent factor
  • λmj\lambda_{mj}: corresponding loading
  • gsig_{si}: study-specific latent factor
  • m=1,,Mm=1,\dots,M0: corresponding study-specific loading
  • m=1,,Mm=1,\dots,M1: study-specific, modality-shared scalar factor
  • m=1,,Mm=1,\dots,M2: overdispersion/random effect term

The latent variables are Gaussian:

m=1,,Mm=1,\dots,M3

and

m=1,,Mm=1,\dots,M4

with m=1,,Mm=1,\dots,M5 diagonal. The overdispersion term is included to capture within-modality correlation and extra variation for non-Gaussian outcomes (Liu et al., 14 Jul 2025).

For the most emphasized modalities, the model treats three variable types:

  • continuous: m=1,,Mm=1,\dots,M6
  • count: m=1,,Mm=1,\dots,M7
  • binary/categorical: m=1,,Mm=1,\dots,M8 is binomial/logistic-style with

m=1,,Mm=1,\dots,M9

This jointly accommodates multi-study structure, multi-modality structure, mixed variable types, covariate effects, and study-shared, study-specific, and modality-specific variation.

3. Latent decomposition, covariates, and identifiability

A central feature of MMGFM is the separation of latent signal into interpretable components. Study-shared factors ss0 represent latent structure common across all studies; their loadings ss1 are modality- and variable-specific but not study-specific. Study-specific factors ss2 represent latent variation unique to study ss3, with loadings ss4 varying by study and modality. The modality-shared random effect ss5 captures correlation among variables within the same modality and study. Covariates enter through the linear term ss6, allowing measured information to explain part of the signal rather than being absorbed into latent factors (Liu et al., 14 Jul 2025).

The paper explicitly states that the model is not identifiable without constraints, and that identifiability conditions are established in Appendix A; it also notes that Condition (C1) in the asymptotic section guarantees identifiability. This is significant because factor models otherwise admit arbitrary rotations, sign changes, and confounding between covariate effects and latent factors. In this setting, identifiability underwrites the interpretation of what is shared across studies, what is study-specific, and how covariates contribute.

In the simulation section, identifiability is also enforced in data generation through SVD-based normalization of loading matrices, for example

ss7

followed by block selection to form shared and specific loadings. This suggests that interpretability is not treated as a purely formal property but as a design constraint affecting both theory and empirical evaluation.

4. Variational approximation and estimation procedure

The observed log-likelihood for observation ss8 in study ss9 is

mm0

where mm1 collects the model parameters in study mm2. The source paper characterizes this likelihood as analytically intractable because it integrates over four large latent random matrices/vectors (Liu et al., 14 Jul 2025).

To address this, it introduces a mean-field variational approximation

mm3

with a fully factorized Gaussian variational family. The corresponding variational lower bound is

mm4

By Jensen’s inequality,

mm5

with equality if and only if the variational density equals the true posterior. The variational posterior therefore acts as the best approximation to the intractable posterior within the chosen mean-field family.

Estimation proceeds via a variational EM algorithm. In the E-step, the variational parameters are updated by

mm6

Because mm7 are not conjugate for the non-Gaussian likelihood, the update uses Laplace approximation combined with Taylor approximation. In the M-step, the profiled variational objective

mm8

is maximized over model parameters. The paper defines the maximum variational lower bound estimator and the maximum variational log-likelihood estimator, and proves that they are equal. It also states that for each mm9, Xsm=(xsimj)ins,  jpmRns×pm,X_{sm}=(x_{simj})_{i\le n_s,\; j\le p_m}\in\mathbb{R}^{n_s\times p_m},0 is the unique maximizer of the lower bound.

5. Asymptotic properties and factor-number selection

The asymptotic theory profiles out the variational parameters and treats the resulting objective as an M-estimation problem. Under conditions Xsm=(xsimj)ins,  jpmRns×pm,X_{sm}=(x_{simj})_{i\le n_s,\; j\le p_m}\in\mathbb{R}^{n_s\times p_m},1–Xsm=(xsimj)ins,  jpmRns×pm,X_{sm}=(x_{simj})_{i\le n_s,\; j\le p_m}\in\mathbb{R}^{n_s\times p_m},2, with Xsm=(xsimj)ins,  jpmRns×pm,X_{sm}=(x_{simj})_{i\le n_s,\; j\le p_m}\in\mathbb{R}^{n_s\times p_m},3 and Xsm=(xsimj)ins,  jpmRns×pm,X_{sm}=(x_{simj})_{i\le n_s,\; j\le p_m}\in\mathbb{R}^{n_s\times p_m},4, Theorem 3 gives the rates

Xsm=(xsimj)ins,  jpmRns×pm,X_{sm}=(x_{simj})_{i\le n_s,\; j\le p_m}\in\mathbb{R}^{n_s\times p_m},5

and

Xsm=(xsimj)ins,  jpmRns×pm,X_{sm}=(x_{simj})_{i\le n_s,\; j\le p_m}\in\mathbb{R}^{n_s\times p_m},6

The paper interprets these rates as reflecting information pooling: shared/modality-wide parameters are estimated using all Xsm=(xsimj)ins,  jpmRns×pm,X_{sm}=(x_{simj})_{i\le n_s,\; j\le p_m}\in\mathbb{R}^{n_s\times p_m},7 samples, whereas study-specific parameters are informed only by Xsm=(xsimj)ins,  jpmRns×pm,X_{sm}=(x_{simj})_{i\le n_s,\; j\le p_m}\in\mathbb{R}^{n_s\times p_m},8 observations from study Xsm=(xsimj)ins,  jpmRns×pm,X_{sm}=(x_{simj})_{i\le n_s,\; j\le p_m}\in\mathbb{R}^{n_s\times p_m},9 (Liu et al., 14 Jul 2025).

Theorem 4 provides an asymptotic linear expansion,

zsiRdz_{si}\in\mathbb{R}^d0

and asymptotic normality for parameter blocks. Examples given in the source include

zsiRdz_{si}\in\mathbb{R}^d1

with analogous results for zsiRdz_{si}\in\mathbb{R}^d2, and for study- and modality-specific parameters,

zsiRdz_{si}\in\mathbb{R}^d3

zsiRdz_{si}\in\mathbb{R}^d4

zsiRdz_{si}\in\mathbb{R}^d5

For dimensionality selection, the paper proposes a step-wise singular value ratio (SVR) criterion. After fitting with generous upper bounds zsiRdz_{si}\in\mathbb{R}^d6 and zsiRdz_{si}\in\mathbb{R}^d7, and writing zsiRdz_{si}\in\mathbb{R}^d8 for the estimated loading matrix for modality zsiRdz_{si}\in\mathbb{R}^d9, the shared factor number is estimated by

xsimjysimjc.i.d.EF(gm(ysimj)),ysimj=τsim+zsiβmj+λmjfi+γsmjgsi+vsim+εsimj.x_{simj}\mid y_{simj} \stackrel{c.i.d.}{\sim} EF(g_m(y_{simj})),\qquad y_{simj} = \tau_{sim} + z_{si}^\top \beta_{mj} + \lambda_{mj}^\top f_i + \gamma_{smj}^\top g_{si} + v_{sim} + \varepsilon_{simj}.0

where xsimjysimjc.i.d.EF(gm(ysimj)),ysimj=τsim+zsiβmj+λmjfi+γsmjgsi+vsim+εsimj.x_{simj}\mid y_{simj} \stackrel{c.i.d.}{\sim} EF(g_m(y_{simj})),\qquad y_{simj} = \tau_{sim} + z_{si}^\top \beta_{mj} + \lambda_{mj}^\top f_i + \gamma_{smj}^\top g_{si} + v_{sim} + \varepsilon_{simj}.1 is the xsimjysimjc.i.d.EF(gm(ysimj)),ysimj=τsim+zsiβmj+λmjfi+γsmjgsi+vsim+εsimj.x_{simj}\mid y_{simj} \stackrel{c.i.d.}{\sim} EF(g_m(y_{simj})),\qquad y_{simj} = \tau_{sim} + z_{si}^\top \beta_{mj} + \lambda_{mj}^\top f_i + \gamma_{smj}^\top g_{si} + v_{sim} + \varepsilon_{simj}.2-th largest singular value. If the xsimjysimjc.i.d.EF(gm(ysimj)),ysimj=τsim+zsiβmj+λmjfi+γsmjgsi+vsim+εsimj.x_{simj}\mid y_{simj} \stackrel{c.i.d.}{\sim} EF(g_m(y_{simj})),\qquad y_{simj} = \tau_{sim} + z_{si}^\top \beta_{mj} + \lambda_{mj}^\top f_i + \gamma_{smj}^\top g_{si} + v_{sim} + \varepsilon_{simj}.3’s differ across modalities, the estimator is

xsimjysimjc.i.d.EF(gm(ysimj)),ysimj=τsim+zsiβmj+λmjfi+γsmjgsi+vsim+εsimj.x_{simj}\mid y_{simj} \stackrel{c.i.d.}{\sim} EF(g_m(y_{simj})),\qquad y_{simj} = \tau_{sim} + z_{si}^\top \beta_{mj} + \lambda_{mj}^\top f_i + \gamma_{smj}^\top g_{si} + v_{sim} + \varepsilon_{simj}.4

otherwise, xsimjysimjc.i.d.EF(gm(ysimj)),ysimj=τsim+zsiβmj+λmjfi+γsmjgsi+vsim+εsimj.x_{simj}\mid y_{simj} \stackrel{c.i.d.}{\sim} EF(g_m(y_{simj})),\qquad y_{simj} = \tau_{sim} + z_{si}^\top \beta_{mj} + \lambda_{mj}^\top f_i + \gamma_{smj}^\top g_{si} + v_{sim} + \varepsilon_{simj}.5 is taken as the mode of xsimjysimjc.i.d.EF(gm(ysimj)),ysimj=τsim+zsiβmj+λmjfi+γsmjgsi+vsim+εsimj.x_{simj}\mid y_{simj} \stackrel{c.i.d.}{\sim} EF(g_m(y_{simj})),\qquad y_{simj} = \tau_{sim} + z_{si}^\top \beta_{mj} + \lambda_{mj}^\top f_i + \gamma_{smj}^\top g_{si} + v_{sim} + \varepsilon_{simj}.6. With the selected xsimjysimjc.i.d.EF(gm(ysimj)),ysimj=τsim+zsiβmj+λmjfi+γsmjgsi+vsim+εsimj.x_{simj}\mid y_{simj} \stackrel{c.i.d.}{\sim} EF(g_m(y_{simj})),\qquad y_{simj} = \tau_{sim} + z_{si}^\top \beta_{mj} + \lambda_{mj}^\top f_i + \gamma_{smj}^\top g_{si} + v_{sim} + \varepsilon_{simj}.7, the model is refit and the same SVR principle is applied to study-specific loading matrices xsimjysimjc.i.d.EF(gm(ysimj)),ysimj=τsim+zsiβmj+λmjfi+γsmjgsi+vsim+εsimj.x_{simj}\mid y_{simj} \stackrel{c.i.d.}{\sim} EF(g_m(y_{simj})),\qquad y_{simj} = \tau_{sim} + z_{si}^\top \beta_{mj} + \lambda_{mj}^\top f_i + \gamma_{smj}^\top g_{si} + v_{sim} + \varepsilon_{simj}.8 to obtain xsimjysimjc.i.d.EF(gm(ysimj)),ysimj=τsim+zsiβmj+λmjfi+γsmjgsi+vsim+εsimj.x_{simj}\mid y_{simj} \stackrel{c.i.d.}{\sim} EF(g_m(y_{simj})),\qquad y_{simj} = \tau_{sim} + z_{si}^\top \beta_{mj} + \lambda_{mj}^\top f_i + \gamma_{smj}^\top g_{si} + v_{sim} + \varepsilon_{simj}.9.

6. Simulation evidence, real-data application, and software

The simulation program is organized into Scenarios 1–4, which likely explains why an informal label such as MFMDScen might arise, although the paper itself does not define that term. In Scenario 1, MMGFM is compared with GFM, MRRR, MSFR, and MultiCOAP in a three-study, three-Poisson-modality setting with covariates and varying fif_i0. The reported findings are that MMGFM consistently outperforms competitors in estimating factors and regression coefficients, remains robust as intramodality correlation fif_i1 increases, and can still estimate loadings well even when within-modality correlation is present, whereas methods that ignore this structure degrade substantially (Liu et al., 14 Jul 2025).

In Scenario 2, which uses mixed Gaussian and Poisson modalities, MMGFM achieves the best mean trace statistics and smallest coefficient error. Scenario 3 varies modality types, sample size, dimension, number of studies, and number of modalities; the summary reported in the source is that MMGFM remains dominant across these settings, while competing methods often break down or perform poorly as the problem becomes more complex. Scenario 4 evaluates factor-number selection and reports that the step-wise SVR method accurately identifies fif_i2 and fif_i3, with accuracy improving as sample size increases and declining as noise or dimension increases.

The real-data application analyzes CITE-seq single-cell multimodal sequencing data from 12 PBMC subjects with COVID-19 status: 4 severe, 3 moderate, and 5 healthy. These three groups are treated as three studies. The two modalities are gene expression counts and CLR-normalized protein markers, and the covariates include age, sex, and days since symptom onset. Using the proposed factor-selection criterion, the fitted dimensions are

fif_i4

The reported outcomes are that MMGFM captures both gene and protein information well, outperforms methods that only handle one modality type or cannot separate study-specific structure, and produces extracted features that support joint clustering and biological interpretation. The study-specific loading matrices are further used to identify potentially important genes and proteins related to immune response and COVID severity; the inferred clusters are said to align with known biology.

The method is implemented in the publicly available R package MMGFM on CRAN. The package provides functionality for fitting the model, performing variational EM estimation, and carrying out factor-number selection via the step-wise SVR rule.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to MFMDScen.