- The paper introduces a Bayesian heterogeneous relative‐shift regression model that fuses coefficient contrasts to reveal collapsed effect patterns.
- It employs a projection‐based generalized ℓ1‐ball prior to stabilize inference and accurately recover latent clusters in high-dimensional settings.
- Applications in macroeconomics and spatial transcriptomics demonstrate improved clustering, reduced MSE, and enhanced interpretability over standard models.
Learning Collapsed Patterns in Compositional Data: Bayesian Heterogeneous Relative-Shift Regression
Introduction
The paper "Learning Collapsed Patterns in Compositional Data: A Bayesian Heterogeneous Relative-Shift Approach" (2606.07373) develops a Bayesian framework for regression with high-dimensional compositional predictors, addressing the dual problems of latent subpopulation heterogeneity and the presence of collapsed (i.e., blockwise-constant) regression effects. By extending the relative-shift regression model to a mixture-of-regressions structure with direct regularization in the contrast space, and by leveraging projection-based priors, the method provides exact contrast-level aggregation and stabilizes high-dimensional posterior inference.
Methodological Framework
The foundation is the relative-shift regression model, which emphasizes identifiable regression contrasts rather than non-identifiable coefficient levels under simplex constraints. For observations yi with compositional covariates xi∈Sp−1, the regression is parameterized so that xi⊤θ enters the mean function, but only pairwise differences θj−θℓ are identifiable. Direct interpretability comes from quantifying the response change under unit mass reallocation between components.
The model extends to the heterogeneous regime by introducing latent cluster indicators zi for a mixture-of-regressions, so the composition--response relationship is allowed to vary across unobserved subpopulations. Cluster-specific coefficients θk inherit the simplex-induced nonidentifiability, preserving the need to regularize contrast space only. In high-dimensional settings, standard unconstrained or coordinatewise-shrinkage priors are shown to cause catastrophic posterior degeneracy: the partition collapses either to complete pooling or maximal separation, making accurate recovery of both clusters and collapsed structures impossible.
The proposed solution is to employ a projection-based generalized ℓ1-ball prior (GLBP), which deterministically projects unconstrained Gaussian precursors to the polyhedral set defined by upper bounds on the sum of contrast absolute values (i.e., total variation in the coefficients). By design, the polytope’s faces correspond to sets of tied coefficients, so the induced prior places positive probability exactly on collapsed (fused) patterns. Each mixture component’s aggregation structure is data-driven, with a separate collapse radius rk∼Exp(ar). The latent partition is modeled by a mixture-of-finite-mixtures (MFM) prior, which is preferred over the Dirichlet process due to its consistency in cluster number estimation.
Posterior Computation
The algorithmic underpinning is a hybrid MCMC scheme embedding structural regularization in the NUTS framework. Discrete cluster labels are marginalized for differentiability, and the projection (surrogate collapse operator) is implemented in closed form by soft-thresholding sorted coefficient differences. Label-switching is handled explicitly by co-clustering matrix alignment, using the Dahl partition as a point estimate with permutation-resolved averaging.
Theoretical Properties
The analysis establishes strong identifiability in the sense of second-order Fisher information separation, ensuring that both the mixing measure and cluster-specific effect structures are statistically learnable. Posterior consistency is proven: the posterior concentrates at optimal rates on both the true number of clusters and the true latent structure, overcoming the high-dimensional degeneracy of unregularized Bayesian mixtures.
Simulation Results
Empirical results on synthetic data demonstrate that the GLBP-based model outperforms baseline mixture-of-regressions with unconstrained or coordinatewise shrinkage priors. In all regimes, including those with p≳n/K, the model shows:
- High adjusted Rand indices (ARI) for partition recovery, often achieving perfect clustering.
- Substantial reductions in coefficient mean squared error (MSE), especially as the ambient dimension p increases.
- Robustness under low sample sizes per cluster, preventing cluster fragmentation due to noise.
Application: Cross-Country Macroeconomic Analysis
The model is applied to 2019 macroeconomic data for 105 countries, regressing log PPP-adjusted GDP per capita on 18 ISIC sectoral shares plus non-compositional controls. The method discovers latent clusters with distinct sector-income relationships, but crucially, sector coefficients within clusters are collapsed to a small number of blocks, demonstrating blockwise-constant effect patterns.

Figure 1: Estimated cluster assignments for the 105 countries under the proposed model, based on the Dahl partition.
Notably, the incremental xi∈Sp−10 from compositional predictors, after accounting for controls, is substantially lower with GLBP regularization compared to an unstructured baseline. This indicates that predictive power in sectors is captured largely by a few interpretable aggregated contrasts. Cluster sizes and average GDP differ only weakly, implying effect patterns are not driven exclusively by income.

Figure 2: Left: posterior co-fusion frequencies for sector pairs across clusters. Right: cluster-specific posterior mean coefficient values.
Within clusters, the effective model dimension is reduced from 18 to between 5 and 9 (Table 1 in the paper), showing strong parsimony without prediction loss. Co-fusion matrices reveal repeated and reproducible collapses of sector coefficients, reflecting interpretable structural aggregation in economic effects.
Application: Spatial Transcriptomics
The method is further validated on spatial transcriptomics from a human lung adenocarcinoma tissue section. For SPP1 gene expression, spot-level cell-type proportions are regressed with Poisson likelihood, revealing strong region-specific aggregation in cell-type effects and spatially contiguous clusters.

Figure 3: Estimated cluster assignments mapped onto the spatial coordinates of the lung adenocarcinoma tissue section. Each spot is colored by its Dahl-partition cluster; spatially contiguous patches of the smaller clusters indicate biologically and geographically structured heterogeneity in the composition--SPP1 relationship.

Figure 4: Fused cell type coefficients xi∈Sp−11 across the SPP1 clusters, with groupings corresponding to identically fused coefficients.
Tumor cells and macrophages universally drive high SPP1 expression, but smaller clusters exhibit region-specific pattern shifts, e.g., ciliated epithelial cell effects emerging only locally. The fitted aggregation fuses large sets of weak or highly correlated cell types, reducing overfitting and enhancing biological interpretability.
Implications and Future Directions
This methodology provides a theoretically sound, computationally efficient mechanism for detection of latent clusters and interpretably collapsed effect structures in high-dimensional compositional regression, even when xi∈Sp−12.
Notable implications include:
- Stable inference under high-dimensionality: Direct regularization on contrasts, rather than on coefficients, is required for robust partition learning in mixture of regressions models.
- Post hoc reliability: Exact coefficient fusion by projection—rather than soft penalties—aligns with interpretable effect identification and credible uncertainty quantification.
- Applicability in genomics and macroeconomics: The model handles domain-specific features such as simplex constraints, high zero inflation, and known or hypothesized structural aggregation (e.g., sector hierarchies, taxonomies).
Potential future extensions involve elaborating the contrast operator to respect problem-specific hierarchies, integrating structured partitions (e.g., Markov random fields for spatial clustering), and accommodating interaction effects or more general dependence structures.
Conclusion
The Bayesian heterogeneous relative-shift regression framework with generalized xi∈Sp−13-ball priors provides a principled solution for learning both latent heterogeneity and collapsed effect patterns in compositional data. The methodology delivers substantive improvements in clustering and parameter estimation over alternatives, supports interpretable effect grouping, and is theoretically justified to avoid degenerate inference regimes that plague high-dimensional unconstrained mixture models. This approach is broadly adaptable to compositional modeling tasks in economics, genomics, and any domain with simplex-structured predictors and unknown latent structure.