Graphical Dirichlet-Type Priors
- Graphical Dirichlet-Type Priors are specialized Bayesian priors that extend classical Dirichlet laws to model complex dependency structures in graphs.
- They utilize clique-separator decompositions and strong hyper-Markov properties to achieve conjugate updating in both discrete and continuous settings.
- Variants such as G-Dirichlet, P-Dirichlet, and Dirichlet–Laplace priors provide flexible frameworks for multinomial modeling, Gaussian inference, and network clustering.
Graphical Dirichlet-type priors constitute a class of Bayesian prior distributions for models with graphical structure, extending the Dirichlet and hyper-Dirichlet laws to accommodate the nuanced conditional independence, directionality, and hierarchical constraints of complex graphs—both directed and undirected, parametric and nonparametric. Their construction exploits the combinatorial and probabilistic structure of decomposable graphs, directed acyclic graphs (DAGs), and block or cluster partitions, yielding conjugacy, strong Markov properties, and flexible hyperparameter regimes suitable for a wide spectrum of applications in graphical modeling and clustering.
1. Fundamental Constructions
The archetypal Dirichlet prior, defined for probability vectors, is generalized to graphical settings by imposing constraints tied to the Markov properties of a dependency graph. In discrete multiway tables governed by a decomposable undirected skeleton , the hyper-Dirichlet and its generalization, the P-Dirichlet law, introduce prior densities over cell probabilities that respect both graphical interconnectedness and, where applicable, restrictions on causal directionality prescribed by a set of moral DAGs sharing as skeleton (Massam et al., 2014). For continuous domains (such as covariance and precision matrices in Gaussian graphical models), analogous graphical Dirichlet-type priors serve both for cell probabilities and structured dependence parameters.
Further expansions include:
- Graphical Dirichlet (G-Dirichlet): Defined by explicit functions of the graph’s clique structure, suitable for modeling multinomial or negative-multinomial cell probabilities over decomposable graphs, and characterized by recursive clique–separator decompositions (Danielewska et al., 2023).
- Dirichlet–Laplace (DL) graphical prior: Employed as a shrinkage prior over Gaussian graphical model precision matrices, utilizing a global-local structure with Dirichlet-weighted local scale parameters (Banerjee, 2019).
- Graphical Dirichlet Process (GDP): Enables clustering with dependency across group-specific random measures characterized via an arbitrary DAG, ensuring the prior over measures respects the conditional independence structure of the underlying graph (Chakrabarti et al., 2023).
- Dirichlet-type priors on directed random graphs: E.g., the infinite relational digraphon model (di-IRM), where Dirichlet priors govern blockwise edge patterns in exchangeable directed random graphs (Cai et al., 2015).
2. Hyperparameterization and Flexibility
The graphical Dirichlet-type family is parametrized by hyperparameters associated with the cliques, separators, and, in the P-Dirichlet case, generalized “clique” and “separator” sets (, ) derived from intersecting the corresponding structures across the admissible graphs in . For the hyper-Dirichlet, the classical constraints
guarantee consistency across overlaps; the P-Dirichlet replaces these with only those linear constraints imposed by partial orderings in , resulting in increased hyperparameter space and, thus, more nuanced control (Massam et al., 2014).
For continuous graphical Dirichlet priors, the shape parameters (e.g., in the G-Dirichlet) involve vector-valued assignments over vertices and control both marginal variances and degrees of neutrality, with specific reductions to the classical Dirichlet or product-of-betas in special cases (complete or edgeless graphs) (Danielewska et al., 2023).
3. Independence Structure and Markov Properties
All graphical Dirichlet-type priors are defined to respect and encode strong (hyper-)Markov properties. In the P-Dirichlet law, for every 0, the conditional probability vectors over a variable given its parent configuration are mutually independent and Dirichlet-distributed. This property, known as the strong directed hyper-Markov property, ensures that posteriors remain within the same family after Bayesian updating—guaranteeing local computational updates and modular inference schemes (Massam et al., 2014).
The G-Dirichlet prior further embodies a graphical neutrality property: for any moral DAG with skeleton 1, the factorizing coordinates (constructed using descendant sets) are mutually independent, each following a beta distribution with parameters determined by the parent set and clique structure. This property yields the strong hyper-Markov structure, allowing updates of marginals/cliques to decompose cleanly (Danielewska et al., 2023).
In Bayesian nonparametric constructions, such as the GDP, the Markov property is ensured through the DAG-based definition of dependency among group-specific Dirichlet processes, whereby each base measure is a mixture over the parent group measures with Dirichlet-distributed mixing weights, and the concentration parameters follow a DAG-hyperprior (Chakrabarti et al., 2023).
4. Characterization and Reduction to Classical Priors
Graphical Dirichlet-type priors unify and generalize the classical Dirichlet and hyper-Dirichlet families through their marginal and dependency structure:
- P-Dirichlet: When 2 is the set of all moral DAGs on 3, the P-Dirichlet reduces to the hyper-Dirichlet. When 4 is complete and 5 consists of separating pairs of orderings, it reduces to the standard Dirichlet (Massam et al., 2014).
- G-Dirichlet: Reduces to standard Dirichlet when 6 is complete (7), and to product-of-betas for an edgeless 8; its dual, the graph-inverted Dirichlet, serves as a conjugate prior for graph multinomial models (Danielewska et al., 2023).
- GDP: Subsumes the hierarchical Dirichlet process (HDP) as a special case (star/fork DAG), and more generally the nested Dirichlet process, but allows intermediate degrees of sharing corresponding to arbitrary DAG structures (Chakrabarti et al., 2023).
- di-IRM: Generalizes the classical Dirichlet mixture model for partitioning observed networks by placing Dirichlet priors on the edge-patterns of each cluster pair, allowing flexible representations of block/relational structure in directed graphs (Cai et al., 2015).
The P-Dirichlet has a crisp characterization via parameter independence: if the joint law of conditional-probability vectors meets mutual independence (global and local) for each graph in a separating family 9, the law must be P-Dirichlet.
5. Conjugacy and Bayesian Updating
Graphical Dirichlet-type priors preserve conjugacy with respect to standard likelihood models arising in graphical inference:
- For multinomial observations under discrete graphical models, the posterior under the (P-)Dirichlet prior is again (P-)Dirichlet with updated hyperparameters by simple additive increments (Massam et al., 2014, Danielewska et al., 2023).
- For graphical negative multinomial and multinomial models, the G-Dirichlet and its inverted analog remain conjugate, with updates following the clique-separator decomposition and strong hyper-Markov property (Danielewska et al., 2023).
- In the DL graphical prior, conjugacy is maintained with respect to the Gaussian graphical model: the latent scale-parameters are updated analytically or by block Gibbs steps, leveraging the Dirichlet coupling for efficient sampling (Banerjee, 2019).
- In the GDP, stick-breaking atoms and cluster assignments are updated via blocked Gibbs samplers, with dependence on the DAG structure respected in all steps via binary tables and logit-transform samplers (Chakrabarti et al., 2023).
6. Structural and Application-Specific Extensions
Graphical Dirichlet-type priors adapt to various structural modeling needs:
- Partial directionality and practitioner constraints: The P-Dirichlet framework assigns zero prior mass to DAGs that violate compulsory directions, supporting both prior elicitation and adherence to domain knowledge (Massam et al., 2014).
- Non-exchangeable group structure: The GDP facilitates clustering where group-specific measures exhibit DAG-governed dependencies and partial atom-sharing—not possible with the HDP or nDP (Chakrabarti et al., 2023).
- Exchangeable network models: In di-IRM, Dirichlet-type priors enable nonparametric block models for exchangeable directed graphs, with closed-form collapsed likelihood integrals (Cai et al., 2015).
- Shrinkage for high-dimensional precision estimation: The DL graphical prior achieves computationally tractable Bayesian inference for sparse precision matrices with near-minimax posterior contraction rates (Banerjee, 2019).
7. Summary of Main Theoretical Properties
| Prior Family | Graph Structure | Conjugacy | Markov Property | Characterization Criterion |
|---|---|---|---|---|
| P-Dirichlet | Discrete DAGs | Yes (Multinomial) | Strong dir. hyper-Markov | Local/global param. independence (Massam et al., 2014) |
| G-Dirichlet | Decomposable G | Yes (Graph. multinom) | Neutrality/Strong hM | Neutrality of factor coordinates (Danielewska et al., 2023) |
| DL-Graphical | Gaussian graphs | Yes (Gaussian MLE) | — | Global-local shrinkage (Banerjee, 2019) |
| GDP | DAG (grouped) | Yes (DP mixtures) | DAG Markov for DPs | Stick, restaurant, hypergraph reps (Chakrabarti et al., 2023) |
| di-IRM | Exchangeable digraphs | Yes (edge clusters) | Blockwise independence | Hierarchical Dirichlet blocks (Cai et al., 2015) |
Each construction is tailored to the independence, directionality, and clustering phenomena intrinsic to the modeled data, while retaining conjugacy, tractable updating, and rigorous interpretative criteria. These priors provide a unifying theme in modern graphical Bayesian inference, bridging discrete, continuous, and nonparametric methodologies, and undergird applications in structure learning, clustering, covariance estimation, and network modeling.