Papers
Topics
Authors
Recent
Search
2000 character limit reached

Modality Missing-Not-at-Random (MMNAR)

Updated 12 July 2026
  • MMNAR is a modality-level missing data pattern where missing modalities are driven by latent causes and systematic clinical decisions rather than random dropout.
  • It employs joint modeling of data and missingness masks, integrating techniques like pattern-recursive imputation and uncertainty-aware fusion for robust multimodal representation.
  • Empirical studies on datasets like MIMIC-IV show that modeling informative missingness leads to significant predictive performance gains in clinical settings.

Modality Missing-Not-at-Random (MMNAR) denotes multimodal settings in which the availability of an entire modality is itself informative: whether a modality is observed or absent is shaped by latent state, acquisition policy, clinician behavior, or the unseen modality content, rather than by random nuisance missingness alone. In recent clinical multimodal work, MMNAR is defined precisely in this sense: modality observation patterns are endogenous, clinically meaningful, and predictive of outcomes, with examples such as discharge summaries missing for 24.5%24.5\% of MIMIC-IV patients and imaging or documentation being ordered selectively by severity, diagnostic complexity, or institutional practice (Liang et al., 21 Sep 2025). More generally, the standard missing-data taxonomy writes

pϕ(MX)={pϕ(M)(MCAR) pϕ(MXobs)(MAR) pϕ(MXobs,Xmis)(MNAR),p_\phi(M\mid X)= \begin{cases} p_\phi(M) & \text{(MCAR)}\ p_\phi(M\mid X^{\text{obs}}) & \text{(MAR)}\ p_\phi(M\mid X^{\text{obs}},X^{\text{mis}}) & \text{(MNAR)}, \end{cases}

and MMNAR is the modality-level instantiation of the third regime (Sim et al., 25 May 2026).

1. Definition and formal scope

A modality-level missingness pattern can be represented by a binary mask M=(M(1),,M(K))M=(M^{(1)},\dots,M^{(K)}), where M(k)M^{(k)} indicates whether modality kk is available. In the MMNAR regime, the probability of a modality pattern need not be a function only of the observed modalities; it may depend on unobserved modality values or on latent causes that jointly influence both modality content and modality presence. This is the statistical core of non-ignorability, carried from classical MNAR theory into multimodal learning (Wang et al., 2021).

Clinical multimodal datasets make this concrete. In the MIMIC-IV setting used for causal representation learning under MMNAR, structured data are available for 100%100\% of patients, discharge summaries for 75.5%75.5\%, radiology reports for 85.0%85.0\%, at least one text modality for 89.5%89.5\%, and chest X-rays for 26.1%26.1\% (Liang et al., 21 Sep 2025). The paper attributes these highly nonuniform observation rates to clinician-assigned observation patterns, institutional protocols, patient severity, adherence, and resource constraints. A plausible implication is that modality presence itself becomes a proxy variable for latent health state and care process.

The literature also distinguishes direct and indirect relevance to MMNAR. Some methods explicitly treat modality observation patterns as first-class variables in the representation or generative model (Liang et al., 21 Sep 2025, Sim et al., 25 May 2026). Others are only adjacent: they are designed for incomplete multimodal prediction and may perform well under simulated “MNAR” settings, but they do not define or estimate a missingness mechanism pϕ(MX)={pϕ(M)(MCAR) pϕ(MXobs)(MAR) pϕ(MXobs,Xmis)(MNAR),p_\phi(M\mid X)= \begin{cases} p_\phi(M) & \text{(MCAR)}\ p_\phi(M\mid X^{\text{obs}}) & \text{(MAR)}\ p_\phi(M\mid X^{\text{obs}},X^{\text{mis}}) & \text{(MNAR)}, \end{cases}0 (Gong et al., 29 Jan 2026).

2. Structural mechanisms behind non-random modality absence

A central structural point is that non-ignorability does not require direct dependence of missingness on the missing variable itself. In hidden Markov models, missingness can be state-dependent: pϕ(MX)={pϕ(M)(MCAR) pϕ(MXobs)(MAR) pϕ(MXobs,Xmis)(MNAR),p_\phi(M\mid X)= \begin{cases} p_\phi(M) & \text{(MCAR)}\ p_\phi(M\mid X^{\text{obs}}) & \text{(MAR)}\ p_\phi(M\mid X^{\text{obs}},X^{\text{mis}}) & \text{(MNAR)}, \end{cases}1 so the same latent state pϕ(MX)={pϕ(M)(MCAR) pϕ(MXobs)(MAR) pϕ(MXobs,Xmis)(MNAR),p_\phi(M\mid X)= \begin{cases} p_\phi(M) & \text{(MCAR)}\ p_\phi(M\mid X^{\text{obs}}) & \text{(MAR)}\ p_\phi(M\mid X^{\text{obs}},X^{\text{mis}}) & \text{(MNAR)}, \end{cases}2 drives both the observation process and the data-generating process (Speekenbrink et al., 2021). The paper shows that this is already enough to make missingness non-ignorable. For MMNAR, the direct analogue is a shared latent representation pϕ(MX)={pϕ(M)(MCAR) pϕ(MXobs)(MAR) pϕ(MXobs,Xmis)(MNAR),p_\phi(M\mid X)= \begin{cases} p_\phi(M) & \text{(MCAR)}\ p_\phi(M\mid X^{\text{obs}}) & \text{(MAR)}\ p_\phi(M\mid X^{\text{obs}},X^{\text{mis}}) & \text{(MNAR)}, \end{cases}3 or pϕ(MX)={pϕ(M)(MCAR) pϕ(MXobs)(MAR) pϕ(MXobs,Xmis)(MNAR),p_\phi(M\mid X)= \begin{cases} p_\phi(M) & \text{(MCAR)}\ p_\phi(M\mid X^{\text{obs}}) & \text{(MAR)}\ p_\phi(M\mid X^{\text{obs}},X^{\text{mis}}) & \text{(MNAR)}, \end{cases}4 that governs both modality contents and modality-presence indicators. This clarifies a common misconception: a modality can be missing-not-at-random because of shared latent causes, even if no explicit self-dependence on the unseen raw modality is written down.

A second structural regime is the “no self-censoring” model. For multivariate binary outcomes with non-monotone missingness, the No Self-Censoring (NSC) assumption is

pϕ(MX)={pϕ(M)(MCAR) pϕ(MXobs)(MAR) pϕ(MXobs,Xmis)(MNAR),p_\phi(M\mid X)= \begin{cases} p_\phi(M) & \text{(MCAR)}\ p_\phi(M\mid X^{\text{obs}}) & \text{(MAR)}\ p_\phi(M\mid X^{\text{obs}},X^{\text{mis}}) & \text{(MNAR)}, \end{cases}5

Under NSC, the missingness of component pϕ(MX)={pϕ(M)(MCAR) pϕ(MXobs)(MAR) pϕ(MXobs,Xmis)(MNAR),p_\phi(M\mid X)= \begin{cases} p_\phi(M) & \text{(MCAR)}\ p_\phi(M\mid X^{\text{obs}}) & \text{(MAR)}\ p_\phi(M\mid X^{\text{obs}},X^{\text{mis}}) & \text{(MNAR)}, \end{cases}6 may depend on other outcomes and other missingness indicators, including outcomes that are themselves missing, but not on its own value after conditioning on the rest (Ren et al., 2023). The paper proves that, for binary outcomes, NSC and MAR intersect only in MCAR. For MMNAR, this is highly informative because it separates self-dependent modality missingness from cross-modality dependence: a modality may be absent because of the rest of the partially observed system without being self-censored by its own latent content.

A third mechanism is latent-class dependence. In model-based clustering with informative missingness, the MNARpϕ(MX)={pϕ(M)(MCAR) pϕ(MXobs)(MAR) pϕ(MXobs,Xmis)(MNAR),p_\phi(M\mid X)= \begin{cases} p_\phi(M) & \text{(MCAR)}\ p_\phi(M\mid X^{\text{obs}}) & \text{(MAR)}\ p_\phi(M\mid X^{\text{obs}},X^{\text{mis}}) & \text{(MNAR)}, \end{cases}7 model assumes missingness depends on class membership: pϕ(MX)={pϕ(M)(MCAR) pϕ(MXobs)(MAR) pϕ(MXobs,Xmis)(MNAR),p_\phi(M\mid X)= \begin{cases} p_\phi(M) & \text{(MCAR)}\ p_\phi(M\mid X^{\text{obs}}) & \text{(MAR)}\ p_\phi(M\mid X^{\text{obs}},X^{\text{mis}}) & \text{(MNAR)}, \end{cases}8 Although pϕ(MX)={pϕ(M)(MCAR) pϕ(MXobs)(MAR) pϕ(MXobs,Xmis)(MNAR),p_\phi(M\mid X)= \begin{cases} p_\phi(M) & \text{(MCAR)}\ p_\phi(M\mid X^{\text{obs}}) & \text{(MAR)}\ p_\phi(M\mid X^{\text{obs}},X^{\text{mis}}) & \text{(MNAR)}, \end{cases}9 is conditionally independent of M=(M(1),,M(K))M=(M^{(1)},\dots,M^{(K)})0 given class M=(M(1),,M(K))M=(M^{(1)},\dots,M^{(K)})1, it remains marginally informative because class membership is latent and tied to the data (Sportisse et al., 2021). This shows that structured missingness patterns can function as latent-subgroup indicators. In modality terms, systematic absence of an image, note, or lab block may identify clinically distinct subpopulations.

3. Modeling paradigms for MMNAR and adjacent MNAR settings

One major paradigm is explicit joint modeling of data and mask. PRDIM writes

M=(M(1),,M(K))M=(M^{(1)},\dots,M^{(K)})2

and treats the missing mask as non-ignorable under MNAR (Sim et al., 25 May 2026). Its “pattern recognizer” M=(M(1),,M(K))M=(M^{(1)},\dots,M^{(K)})3 approximates M=(M(1),,M(K))M=(M^{(1)},\dots,M^{(K)})4, and the mask model enters both the ELBO and the reverse diffusion updates through a guidance term. This is a direct operationalization of the MMNAR principle that an absence pattern is informative and should constrain imputation or generation.

A related generative perspective appears in GNR, whose abstract states that the complete data and missing mask are treated as “two modalities of incomplete data on an equal footing,” using a “conjunction model” decomposition and concurrent reconstruction of data and mask (Chen et al., 2023). Although the detailed method description is not available in the supplied material, the framing is directly relevant to MMNAR: the availability pattern is modeled as an information-bearing object rather than a passive bookkeeping device.

Another paradigm is pattern-mixture recursion under sparse pattern support. MISPR develops a constructive characterization of the full law by recursively identifying pattern-specific extrapolation densities: M=(M(1),,M(K))M=(M^{(1)},\dots,M^{(K)})5 It then uses a pattern DAG to determine which supported patterns can supply the Gibbs factors needed to impute a given unsupported or partially supported pattern (Phung et al., 21 Jul 2025). This is particularly relevant to MMNAR because unsupported combinations of observed and missing modalities are routine in high-dimensional multimodal systems. A plausible implication is that MMNAR handling should be pattern-sensitive rather than based on indiscriminate borrowing across all modality configurations.

A more indirect but practically important line uses classical selection models to derive modified imputation conditionals. In sequential regression multiple imputation under cross-variable MNAR, the ideal conditional takes the form

M=(M(1),,M(K))M=(M^{(1)},\dots,M^{(K)})6

which leads, after approximation, to imputation models containing missingness indicators, interactions, or fixed offsets derived from explicit missingness models (Beesley et al., 2021). For MMNAR, this provides a classical rationale for feeding modality-presence indicators into modality-imputation models rather than conditioning only on observed content.

4. MMNAR-aware multimodal representation learning

The most explicit multimodal MMNAR framework in the supplied literature is the causal representation learning model for clinical records. For patient M=(M(1),,M(K))M=(M^{(1)},\dots,M^{(K)})7, it defines a modality set M=(M(1),,M(K))M=(M^{(1)},\dots,M^{(K)})8, observed modalities M=(M(1),,M(K))M=(M^{(1)},\dots,M^{(K)})9, observed inputs M(k)M^{(k)}0, and a modality indicator vector M(k)M^{(k)}1 (Liang et al., 21 Sep 2025). The method learns a missingness embedding

M(k)M^{(k)}2

uses it to gate modality embeddings,

M(k)M^{(k)}3

and then performs attention-based multimodal fusion. The observation pattern is therefore embedded, reconstructed, and used to modulate the contribution of each available modality.

This MMNAR-aware fusion is paired with cross-modal reconstruction and contrastive learning. During training, an observed modality is randomly masked, a partial representation M(k)M^{(k)}4 is formed from the remaining modalities, and a decoder reconstructs the held-out modality embedding: M(k)M^{(k)}5 The reconstruction and InfoNCE-style contrastive losses are intended to enforce “semantic sufficiency,” so that the representation captures patient state rather than collapsing into a code for the missingness pattern alone (Liang et al., 21 Sep 2025). This suggests an architectural principle for MMNAR: informative missingness should be modeled, but the learned state should still be able to support cross-modal inference under counterfactual modality configurations.

The same framework adds a rectifier for pattern-specific residual bias. After learning a shared representation M(k)M^{(k)}6, task predictions are decomposed as

M(k)M^{(k)}7

where M(k)M^{(k)}8 is estimated by cross-fitted residual averages within each missingness pattern and used only when its magnitude exceeds a threshold M(k)M^{(k)}9 (Liang et al., 21 Sep 2025). This is not a formal causal identification result, but it is a direct recognition that even a pattern-aware representation may leave residual bias associated with specific clinician-assigned observation patterns.

An adjacent but distinct strategy is uncertainty-aware fusion. AUM models each unimodal embedding as

kk0

uses uncertainty-aware attention with weights inversely related to aggregated uncertainty, and propagates means and variances through a bipartite patient–modality graph (Gong et al., 29 Jan 2026). Missing modalities are conceptually initialized with high uncertainty and receive negligible attention. The paper is explicit, however, that this is not a causal missingness model; it models predictive uncertainty conditional on observed modalities rather than kk1. This suggests a useful distinction within the MMNAR literature: mechanism-aware models use missingness patterns as explicit variables, whereas proxy methods attenuate the damage of non-random missingness by downweighting unreliable or effectively absent evidence.

5. Testing, sensitivity analysis, and design for MMNAR investigation

Because MNAR and MMNAR are generally not identifiable from the primary incomplete sample alone, several papers focus on diagnosis rather than direct correction. One route is null-only testing. Under a logistic missingness model

kk2

testing MAR versus MNAR reduces to

kk3

The score tests of Wang, Lu, and Liu require estimation only under the null MAR model, which circumvents the non-identifiability of the full MNAR alternative (Wang et al., 2021). For MMNAR, this suggests testing whether a modality-presence indicator remains associated with a low-dimensional summary of the missing modality after conditioning on observed modalities.

A second route is targeted recovery or follow-up sampling. In the selection-model framework, a deliberately designed recovery sample makes MNAR testable because it reveals the missing value among cases that were initially unobserved (Noonan et al., 2022). For logistic missingness models, the paper shows that targeted follow-up changes only the intercept of the augmented missingness model, leaving the coefficients governing dependence on the missing value unchanged. The direct MMNAR analogue is targeted modality reacquisition: reacquiring a missing modality for a strategically chosen subset can support valid tests of whether modality absence depends on its own unseen content.

A third route is sensitivity analysis around restricted MNAR models. Under NSC, self-censoring interactions are excluded, but they can be reintroduced via parameters kk4, yielding the shifted conditional model

kk5

These kk6 are interpreted as conditional log-odds ratios linking missingness to the variable’s own value and serve as sensitivity parameters rather than identified quantities (Ren et al., 2023). For MMNAR, this provides a principled decomposition between cross-modality missingness dependence and self-modality missingness dependence.

Finally, modified chained-equations imputation provides a practical but assumption-driven diagnostic layer. The derivations for SRMI under cross-variable MNAR show why missingness indicators, interactions between missingness indicators and observed covariates, and fixed offsets constructed from missingness models should appear in the imputation regression (Beesley et al., 2021). In modality terms, this explains why imputation models that omit the modality mask are misspecified under MMNAR-like cross-modality dependence.

6. Empirical evidence, limitations, and unresolved directions

Direct empirical evidence for MMNAR-aware multimodal prediction is strongest in clinical representation learning. On MIMIC-IV, the MMNAR-aware framework achieves AUC kk7 for 30-day readmission, kk8 for post-discharge ICU admission, and kk9 for in-hospital mortality, with reported relative gains of 100%100\%0, 100%100\%1, and 100%100\%2 over the strongest baselines; on eICU it reaches AUC 100%100\%3 for readmission and 100%100\%4 for mortality, with a reported 100%100\%5 gain for readmission (Liang et al., 21 Sep 2025). The ablations are especially informative: the “+ MMNAR” component alone contributes large improvements before reconstruction and rectification are added, indicating that pattern-aware fusion is not merely a minor auxiliary feature.

Evidence from adjacent work is more qualified. AUM reports improvements of 100%100\%6 AUC-ROC on MIMIC-IV mortality prediction and 100%100\%7 on eICU, and it remains robust up to a missing ratio of 100%100\%8 under a simulated MNAR scenario while maintaining around 100%100\%9 AUC-ROC (Gong et al., 29 Jan 2026). Yet the same paper states that it does not model the missingness process itself. Its relevance to MMNAR is therefore practical rather than formal: heteroscedastic uncertainty and reliability-aware graph aggregation can mitigate some downstream harms of informative missingness without addressing selection bias in the Rubin/Pearl sense.

Generative MNAR methods also show that explicit mask modeling matters outside multimodal fusion proper. PRDIM reports strong MNAR imputation performance across time series, images, and tabular data, with its main mechanism being joint modeling of data and mask via a pattern recognizer (Sim et al., 25 May 2026). MISPR is comparable to MICE under MAR, superior and less biased under MNAR, and specifically designed for sparse pattern support, where unsupported combinations of missingness patterns would otherwise force unjustified extrapolation (Phung et al., 21 Jul 2025). These results suggest that MMNAR research should treat the combinatorics of modality patterns as a primary modeling problem rather than a preprocessing nuisance.

The main limitation across the literature is that explicit MMNAR handling remains only partially solved. Direct multimodal MMNAR work is concentrated in clinical prediction and still relies on proxy causal assumptions, pattern embeddings, and post-hoc rectification rather than a fully identified modality-level selection model (Liang et al., 21 Sep 2025). Adjacent uncertainty-aware multimodal fusion improves robustness but leaves the missingness mechanism unmodeled (Gong et al., 29 Jan 2026). Pattern-recursive and MMD-based estimators provide principled robustness or identification templates, but they are developed for classical variable-level settings rather than high-dimensional multimodal encoders and decoders (Phung et al., 21 Jul 2025, Chérief-Abdellatif et al., 1 Mar 2025). A persistent gap, therefore, is explicit modeling of 75.5%75.5\%0 for high-dimensional modality blocks under realistic sparse-support regimes, together with diagnostics and sensitivity analyses that remain valid when modality absence is driven by latent semantics, acquisition policy, and institutional workflow rather than by random dropout alone.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Modality Missing-Not-at-Random (MMNAR).