Markov Missing Graph (MMG) Framework
- MMG is an undirected-graph framework that decomposes missing data imputation into local conditional models based on graph neighborhoods.
- It employs the Principle of Available Information (PAI) to identify each local submodel using all cases where relevant variables are observed.
- MMG demonstrates improved efficiency and lower bias compared to traditional methods under MCAR and MAR scenarios.
to=arxiv_search 手机版天天中彩票_code
{"query":"id:2509.03410 OR ti:\"Markov Missing Graph: A Graphical Approach for Missing Data Imputation\"","max_results":5,"sort_by":"submittedDate","sort_order":"descending"}【อ่านข้อความเต็มjson
[{"arxiv_id":"2509.03410","title":"Markov Missing Graph: A Graphical Approach for Missing Data Imputation","authors":["Jingyang Wang","Weijie Su","Jin Liu"],"abstract":"We introduce the Markov missing graph (MMG), a novel framework that imputes missing data based on undirected graphs. MMG leverages conditional independence relationships to locally decompose the imputation model. To establish the identification, we introduce the Principle of Available Information (PAI), which guides the use of all relevant observed data. We then propose a flexible statistical learning paradigm, MMG Imputation Risk Minimization under PAI, that frames the imputation task as an empirical risk minimization problem. This framework is adaptable to various modeling choices. We develop theories of MMG, including the connection between MMG and Little's complete-case missing value assumption, recovery under missing completely at random, efficiency theory, and graph-related properties. We show the validity of our method with simulation studies and illustrate its application with a real-world Alzheimer's data set.","categories":["stat.ME","cs.LG","stat.ML"],"published":"2025-09-03","pdf_url":"http://arxiv.org/pdf/2509.03410v1","abs_url":"https://arxiv.org/abs/2509.03410v1"}]
to=arxiv_search 大发快三豹子_code
{"query":"id:2602.08407 OR ti:\"Drop the mask! GAMM-A Taxonomy for Graph Attributes Missing Mechanisms\"","max_results":5,"sort_by":"submittedDate","sort_order":"descending"}【อ่านข้อความเต็มjson
[{"arxiv_id":"2602.08407","title":"Drop the mask! GAMM-A Taxonomy for Graph Attributes Missing Mechanisms","authors":["David Limmer","Nils M. Kriege","Michael T. Schaub"],"abstract":"Exploring missing data in attributed graphs introduces unique challenges beyond those found in tabular datasets. In this work, we extend the taxonomy for missing data mechanisms to attributed graphs by proposing GAMM (Graph Attributes Missing Mechanisms), a framework that systematically links missingness probability to both node attributes and the underlying graph structure. Our taxonomy enriches the conventional definitions of masking mechanisms by introducing graph-specific dependencies. We empirically demonstrate that state-of-the-art imputation methods, while effective on traditional masks, significantly struggle when confronted with these more realistic graph-aware missingness scenarios.","categories":["cs.LG","stat.ML"],"published":"2026-02-09","pdf_url":"http://arxiv.org/pdf/2602.08407v1","abs_url":"https://arxiv.org/abs/2602.08407v1"}]
to=arxiv_search 大发棋牌_code
{"query":"id:1301.6624 OR ti:\"Markovian acyclic directed mixed graphs for discrete data\"","max_results":5,"sort_by":"submittedDate","sort_order":"descending"}【อ่านข้อความเต็มjson
[{"arxiv_id":"1301.6624","title":"Markovian acyclic directed mixed graphs for discrete data","authors":["Robin J. Evans","Thomas S. Richardson"],"abstract":"Acyclic directed mixed graphs (ADMGs) are graphs that contain directed ($\rightarrow$) and bidirected ($\leftrightarrow$) edges, subject to the constraint that there are no cycles of directed edges. Such graphs may be used to represent the conditional independence structure induced by a DAG model containing hidden variables on its observed margin. The Markovian model associated with an ADMG is simply the set of distributions obeying the global Markov property, given via a simple path criterion (m-separation). We first present a factorization criterion characterizing the Markovian model that generalizes the well-known recursive factorization for DAGs. For the case of finite discrete random variables, we also provide a parameterization of the model in terms of simple conditional probabilities, and characterize its variation dependence. We show that the induced models are smooth. Consequently, Markovian ADMG models for discrete variables are curved exponential families of distributions.","categories":["math.ST","stat.TH"],"published":"2013-01-28","pdf_url":"http://arxiv.org/pdf/1301.6624v1","abs_url":"https://arxiv.org/abs/1301.6624v1"}]
to=arxiv_search ทะเบียนฟรี_code
{"query":"id:2109.02770 OR ti:\"Ignorable and non-ignorable missing data in hidden Markov models\"","max_results":5,"sort_by":"submittedDate","sort_order":"descending"}【อ่านข้อความเต็มjson
[{"arxiv_id":"2109.02770","title":"Ignorable and non-ignorable missing data in hidden Markov models","authors":["J. M. Michelot","R. Langrock","T. A. Patterson"],"abstract":"We consider missing data in the context of hidden Markov models with a focus on situations where data is missing not at random (MNAR) and missingness depends on the identity of the hidden states. In simulations, we show that including a submodel for state-dependent missingness reduces bias when data is MNAR and state-dependent, whilst not reducing accuracy when data is missing at random (MAR). When missingness depends on time but not the hidden states, a model which only allows for state-dependent missingness is biased, whilst a model that allows for both state- and time-dependent missingness is not. Overall, these results show that modelling missingness as state-dependent, and including other relevant covariates, is a useful strategy in applications of hidden Markov models to time-series with missing data.","categories":["stat.ME","stat.AP"],"published":"2021-09-06","pdf_url":"http://arxiv.org/pdf/2109.02770v1","abs_url":"https://arxiv.org/abs/2109.02770v1"}]
to=arxiv_search 출장안마_code
{"query":"id:1107.3036 OR ti:\"A note on global Markov properties for mixed graphs\"","max_results":5,"sort_by":"submittedDate","sort_order":"descending"}【อ่านข้อความเต็มjson
[{"arxiv_id":"1107.3036","title":"A note on global Markov properties for mixed graphs","authors":["Steffen Lauritzen","Nanny Wermuth","Vincenzo Caputo"],"abstract":"Global Markov properties in mixed graphs are usually formulated in terms of the path-oriented m-separation or by use of augmented graphs (similar to moral graphs in the case of directed acyclic graphs). We provide an alternative characterization that can be easily implemented.","categories":["math.ST","stat.TH"],"published":"2011-07-15","pdf_url":"http://arxiv.org/pdf/1107.3036v1","abs_url":"https://arxiv.org/abs/1107.3036v1"}]
to=arxiv_search 天天中彩票不中返_code
{"query":"id:2310.16401 OR ti:\"Graph Neural Networks with a Distribution of Parametrized Graphs\"","max_results":5,"sort_by":"submittedDate","sort_order":"descending"}【อ่านข้อความเต็มjson
[{"arxiv_id":"2310.16401","title":"Graph Neural Networks with a Distribution of Parametrized Graphs","authors":["Mingxuan Ju","Alicia Caro Yustas","Sean X. Wang","Mingyi Hong","Lorenzo Torresani","Mladen Kolar"],"abstract":"Traditionally, graph neural networks have been trained using a single observed graph. However, the observed graph represents only one possible realization. In many applications, the graph may encounter uncertainties, such as having erroneous or missing edges, as well as edge weights that provide little informative value. To address these challenges and capture additional information previously absent in the observed graph, we introduce latent variables to parameterize and generate multiple graphs. We obtain the maximum likelihood estimate of the network parameters in an Expectation-Maximization (EM) framework based on the multiple graphs. Specifically, we iteratively determine the distribution of the graphs using a Markov Chain Monte Carlo (MCMC) method, incorporating the principles of PAC-Bayesian theory. Numerical experiments demonstrate improvements in performance against baseline models on node classification for heterogeneous graphs and graph regression on chemistry datasets.","categories":["cs.LG","stat.ML"],"published":"2023-10-25","pdf_url":"http://arxiv.org/pdf/2310.16401v1","abs_url":"https://arxiv.org/abs/2310.16401v1"}]
Markov Missing Graph (MMG) is an undirected-graph-based framework for missing-data imputation that uses conditional independence relationships to locally decompose the imputation model. In its basic formulation, MMG starts from a pattern-mixture decomposition and treats the extrapolation density as the primary imputation object; the graph then imposes a local Markov structure under which each connected missing block is imputed only from its graph neighborhood. Identification is supplied by the Principle of Available Information (PAI), which prescribes using all cases in which the variables required by a local submodel are observed, and the resulting framework is developed as both an identification theory and a flexible empirical risk minimization paradigm [2509.03410].
1. Formal definition and graphical factorization
Let (X=(X_1,\dots,X_d)) denote the variables subject to missingness and let (R=(R_1,\dots,R_d)\in{0,1}d) be the response pattern, where (R_j=1) means (X_j) is observed and (R_j=0) means missing. For a given pattern (R=r), the observed variables are (X_r) and the missing variables are (X_{\bar r}). The starting point is the pattern-mixture decomposition
[
p(x,r)=p(x_{\bar r}\mid x_r,R=r)\,p(x_r,R=r),
]
where the key imputation target is the extrapolation density (p(x_{\bar r}\mid x_r,R=r)) [2509.03410].
MMG introduces an undirected graph (G=(V,E)) with vertices (V={X_1,\dots,X_d}), where edges encode conditional dependence among study variables. The graph is used to impose a local Markov structure on imputation: missing variables are imputed using only their graph neighbors, not all observed variables. If the missing nodes decompose into connected components
[
\bar r = s_1+\cdots+s_K,
]
then MMG assumes
[
p(x_{\bar r}\mid x_r,R=r) = \prod_{k=1}K p(x_{s_k}\mid x_{N_G(s_k)},R=r). \tag{MMG}
]
Here (s_k) is a connected component of the missing variables and (N_G(s_k)) is the neighborhood of that component in (G). This factorization is the core structural assumption of MMG. It converts a potentially high-dimensional extrapolation model for all missing entries into a product of localized conditional models, one for each connected missing block.
The significance of this construction is not merely computational. The decomposition is presented as a graph-induced statement about conditional dependence in the imputation model itself. In that sense, MMG is a missing-data model organized around an undirected conditional-independence graph rather than around a single global extrapolation density.
2. Principle of Available Information and identification
The MMG factorization alone does not identify the local conditionals
[
p(x_{s_k}\mid x_{N_G(s_k)},R=r),
]
because (X_{s_k}) is missing under pattern (r). The Principle of Available Information resolves this by specifying which incomplete records may be used to estimate each local submodel. The relevant variable set for a component (s_k) is
[
\bar N_G(s_k)=s_k\cup N_G(s_k),
]
and PAI states that the local imputation model is identified from cases with
[
R_{\bar N_G(s_k)}=1.
]
The identifying restriction is
[
p(x_{s_k}\mid x_{N_G(s_k)},R=r) = p(x_{s_k}\mid x_{N_G(s_k)},R_{\bar N_G(s_k)}=1). \tag{PAI}
]
Under this restriction, estimation of a local MMG submodel uses all observations in which the component and its neighbors are observed, regardless of the global response pattern [2509.03410].
PAI therefore plays two roles. First, it gives a precise identification rule for each local conditional distribution. Second, it defines the admissible information set for estimation. The paper describes this as a “minimal sufficient information” principle for imputation: do not use more than necessary, but do use everything available that is relevant to the local model.
A central theorem states that MMG together with PAI nonparametrically identifies the full-data distribution (p(x,r)). The proof strategy is to combine the pattern-mixture decomposition, the MMG factorization of the extrapolation density, and the PAI replacement of each nonidentifiable local term by a conditional distribution estimable from sufficiently observed cases. This makes MMG+PAI an identification strategy rather than only a modeling convenience.
3. MMG Imputation Risk Minimization under PAI
To turn MMG into a learnable statistical procedure, the framework introduces MMG Imputation Risk Minimization under PAI. The collection of local imputation models is
[
\mathcal{M} = \left{ p(x_s\mid x_{N_G(s)},R=r): s\le \bar r,\; s \text{ connected component of }G,\; r\in{0,1}d \right}.
]
Under PAI, if two response patterns (r_1,r_2) induce the same connected pattern (s), then
[
p(x_s\mid x_{N_G(s)},R=r_1) = p(x_s\mid x_{N_G(s)},R=r_2),
]
so the model can be indexed by (s) alone, written (\theta_{r,s}\equiv\theta_s) [2509.03410].
For a loss (\mathcal{L}(\theta_s\mid x_s,x_{N_G(s)})), the population MMG imputation risk is
[
\mathcal{R}(\theta_s) = \mathbb{E}\left[ \mathcal{L}(\theta_s\mid x_s,x_{N_G(s)}) \cdot I(R_{\bar N_G(s)}=1) \right]. \tag{Risk}
]
The target parameter is
[
\theta_s* = \arg\min_{\theta_s\in\Theta} \mathbb{E}\left[ \mathcal{L}(\theta_s\mid x_s,x_{N_G(s)}) \cdot I(R_{\bar N_G(s)}=1) \right].
]
Given data, the empirical risk is
[
\hat{\mathcal{R}}n(\theta_s) = \frac{1}{n}\sum{i=1}n \mathcal{L}(\theta_s\mid X_{i,s},X_{i,N_G(s)}) \cdot I(R_{i,\bar N_G(s)}=1), \tag{ERM}
]
and the estimator is
[
\hat\theta_s = \arg\min_{\theta_s\in\Theta}\hat{\mathcal{R}}_n(\theta_s). \tag{Estimator}
]
This formulation makes MMG adaptable to multiple model classes. A likelihood-based choice uses negative log-likelihood,
[
\mathcal{L}(\theta_s\mid x_s,x_{N_G(s)}) = -\log p(x_s,x_{N_G(s)};\theta_s),
]
in which case (\hat\theta_s) is the maximum-likelihood estimator on the subset of cases with (R_{i,\bar N_G(s)}=1). The framework is explicitly described as flexible and compatible with different statistical learning choices.
4. Theoretical properties
Theoretical development in MMG covers identification, equivalence results, recovery guarantees, efficiency, and graph-specific structural properties [2509.03410].
A central relation is to Little’s complete-case missing value assumption (CCMV),
[
p(X_{\bar r}\mid X_r,R=r)=p(X_{\bar r}\mid X_r,R=1).
]
If (G) is fully connected, then MMG+PAI is equivalent to CCMV. Under monotone missingness and a chain graph (X_1-X_2-\cdots-X_d) faithful to the complete-case distribution, CCMV and MMG+PAI are also equivalent. In this sense, CCMV appears as a special case of MMG under strong graph-connectivity conditions.
For recovery guarantees, the framework proves that if missingness is MCAR and the distribution of (X) is faithful to graph (G), then NP-MMG with respect to (G) recovers the true model. If (X) is multivariate normal, G-MMG is asymptotically more efficient than complete-case analysis. This connects the graphical decomposition to both correctness and efficiency when the graph is faithful and missingness is completely random.
The efficiency theory is developed for estimating (\mu=\mathbb{E}[X_1]). For patterns whose missing component containing (X_1) is (s_j), the paper defines the regression function
[
m_{s_j}(x_{N_G(s_j)})=
\int x_1\,p(x_1\mid x_{N_G(s_j)},R_{\bar N_G(s_j)}=1)\,dx_1,
]
and the odds-like function
[
\mathcal{O}{s_j}(x{N_G(s_j)})=
\frac{P(R\in\mathbb{S}j\mid X{N_G(s_j)})}
{P(R\ge \bar N_G(s_j)\mid X_{N_G(s_j)})}.
]
From these, the paper derives both a regression adjustment estimator and an inverse probability weighting estimator. The linear part of the efficient influence function for (\mu_{s_j}=\mathbb{E}[X_1I(\psi(\bar R)=s_j)]) is
[
\mathrm{LEIF}{s_j}(X_1,X{N_G(s_j)},R) = X_1\,\mathcal{O}{s_j}(X{N_G(s_j)})\,I(R\ge \bar N_G(s_j)) + m_{s_j}(X_{N_G(s_j)}) \Big( I(\psi(\bar R)=s_j) - I(R\ge \bar N_G(s_j))\,\mathcal{O}{s_j}(X{N_G(s_j)}) \Big).
]
The resulting estimator is multiply robust: under assumptions (A1)–(A4), if either the odds model or the regression model is consistent, the estimator remains consistent.
The assumptions used for these results are: (A1) Donsker/stability, requiring nuisance estimators to lie in uniformly bounded Donsker classes and converge in (L_2(P)); (A2) uniform boundedness of (X_1), (\mathcal{O}{s_j}), and (m{s_j}); (A3) multiply-robustness, meaning that for each (j), at least one nuisance estimator converges to the correct target; and (A4) a uniform entropy bound ensuring that the “use data twice” procedure remains asymptotically valid.
The graph-specific propositions further formalize locality. Proposition 1 establishes local invariance: if two response patterns (r_1,r_2) have the same missing connected component (s) and the same observed neighborhood (N_G(s)), then the same local conditional model applies. Proposition 2 establishes nested submodels: if (B\subset A) and (\bar N_G(A)=\bar N_G(B)), then the larger-set imputation model implies the smaller-set one by marginalization. Proposition 3 establishes locality with respect to graph neighborhoods: if (\bar r) is a single connected component and two graphs (G_1,G_2) have the same neighborhood for (\bar r), then they yield the same imputation model for that pattern. These results make explicit that MMG is locally determined by the neighborhood around the missing component.
5. Implementation, model classes, and workflow
The practical workflow begins with graph specification. If prior knowledge is lacking, the graph (G) may be estimated from complete cases, for example by graphical lasso or sparse neighborhood regression. The paper emphasizes that under MCAR and graph faithfulness, such an estimated graph can be valid for MMG [2509.03410].
Implementation then proceeds through a localized modeling pipeline:
- Choose or estimate the graph: specify (G) from substantive knowledge or estimate it from complete cases.
- Identify connected patterns and model patterns: group response patterns according to connected components of the missing set and the required observed neighborhood (\bar N_G(s)).
- Fit local submodels: instantiate the conditional models using a data-appropriate family.
- Impute from local conditional distributions: for each incomplete record, determine the relevant missing block and draw or predict from the corresponding conditional model.
- Repeat imputation: the paper uses multiple imputations, for example (m=20) completed datasets.
Three concrete MMG variants are described. G-MMG uses Gaussian local models for continuous data. I-MMG uses Ising models for binary data. MP-MMG uses mixture-of-product models for mixed types. The authors also mention an R package, mmg, and note that for the NACC application they adapted code from the mixturebpe package for MP-MMG.
This implementation pattern reflects the architecture of the theory: graph structure determines the locality of each subproblem, PAI determines which observations are admissible for estimation, and the chosen statistical family determines how each local conditional is fitted.
6. Simulation studies and the Alzheimer’s application
The simulation studies use (n=2000), 100 Monte Carlo replicates, and comparisons against complete-case analysis, MICE, missForest, and MMG variants [2509.03410].
For G-MMG, data are generated from Gaussian graphical models with MCAR and MAR missingness. Under MCAR, all methods give comparable median estimates, but complete-case analysis has the highest variability. Under MAR, G-MMG shows low bias, low variability, better performance than MICE and missForest, and substantial efficiency gains over complete-case analysis. For MP-MMG, data are generated from a mixture of two Gaussian graphical models. The reported result is robustness under both MCAR and MAR, with low bias and variability even when the data-generating mechanism is more complex and not exactly a single graphical Gaussian model. In these harder settings, complete-case analysis, MICE, and missForest show more bias or instability.
The principal real-data application uses the National Alzheimer’s Coordinating Center (NACC) data. The study includes (n=27{,}070) individuals, 351 missingness patterns, and only 5,739 complete cases. The application uses demographic variables as fully observed predictors, eight neuropsychological tests, and the Clinical Dementia Rating as an external outcome. The graph is estimated from complete-case test data using graphical lasso with EBIC tuning, reducing 351 patterns to 84 model patterns under MMG. The fitted model is MP-MMG with EM estimation, and the analysis generates 20 multiply imputed datasets.
In the downstream analysis, logistic regression is used to model whether CDR worsens from year 1 to year 2. Relative to complete-case analysis, MMG yields narrower confidence intervals; complete-case analysis sometimes produces coefficient signs that differ from MMG, MICE, and missForest; and MMG gives more stable and clinically plausible results. The reported predictors of higher progression risk include lower baseline performance on global cognition, language, and memory tests, as well as slower processing speed and poorer executive function; worsening TRAILA/TRAILB over time is also associated with progression.
The age-stratified analysis reports that MMG produces smooth, clinically plausible decline patterns across age groups, whereas complete-case analysis shows irregular or counterintuitive patterns, likely due to selective dropout. A sensitivity analysis varying the graph sparsity threshold in graphical lasso finds that estimates are stable for moderate thresholds, although overly sparse graphs can alter some coefficients.
7. Relation to adjacent frameworks and terminological boundaries
The term Markov Missing Graph has a specific meaning in the undirected-graph imputation framework above and should be distinguished from several neighboring uses of “Markov,” “graph,” and “missingness.”
A first source of confusion is the literature on Markovian acyclic directed mixed graphs. Markovian ADMGs are mixed graphs with directed and bidirected edges used to represent the conditional independence structure induced by DAGs after marginalizing hidden variables; their Markov models are defined via m-separation and characterized by head–tail factorizations over ancestral sets. They are not “missing data graphs” in the sense of modeling response indicators or extrapolation densities [1301.6624]. Closely related work on global Markov properties for mixed graphs uses ancestral reduction and augmentation to convert m-separation questions into ordinary undirected separation problems; this is again a conditional-independence construction rather than an MMG imputation framework [1107.3036].
A second neighboring area is graph-aware missingness in attributed graphs. GAMM formalizes an attributed graph as (G=(\mathcal{V},\mathcal{E},F)) with a binary mask (\Omega\in{0,1}{n\times d}) and models missingness through
[
P(\Omega_{ij}=0\mid G)=g(\cdot),
]
where the mechanism may depend on node attributes, structural properties, neighborhood attributes, or combinations thereof. The taxonomy includes MCAR, A-MAR/A-MNAR, S-MAR/S-MNAR, N-MAR/N-MNAR, and G-MAR/G-MNAR, and it explicitly extends traditional missingness definitions by incorporating the interplay between node attributes and graph structure [2602.08407]. A plausible implication is that GAMM and MMG address different layers of the same broad problem: GAMM classifies how missingness may be generated on graphs, whereas MMG specifies how imputation may be locally decomposed once a graph over variables is available.
A third adjacent area is hidden Markov modeling with missing data. In one line of work, a hidden Markov model is extended with a missingness submodel in which the latent state (S_t) can drive both the observation (Y_t) and the missingness indicator (M_t), and time may also enter the missingness mechanism; this treats missingness as non-ignorable when it depends on unobserved states [2109.02770]. In another line, an infinite hidden Markov model for multiple multivariate time series with missing data imputes missing values from the posterior predictive distribution conditional on the hidden state and emission parameters, rather than introducing a separate graph of missingness dependencies [2204.06610]. These models are Markovian and missing-data aware, but they are not MMG in the undirected local-decomposition sense.
Finally, latent-graph learning methods that treat the observed graph as one realization from a distribution over plausible graphs are MMG-adjacent only in a loose sense. For example, a GNN framework based on a distribution of parametrized graphs uses EM, MCMC, and PAC-Bayesian posteriors to learn over uncertain graphs with missing or spurious edges, but it does not define a missing-edge generative model in MMG terms and is best understood as task-driven latent graph distribution learning rather than as a Markov Missing Graph model [2310.16401].
Within this landscape, MMG is most precisely characterized as a graphical, locally decomposed missing-data imputation framework built on two linked principles: a Markov graph structure that restricts each missing block to its graph neighborhood, and PAI, which identifies each local conditional from all cases where the relevant variables are observed [2509.03410].