Papers
Topics
Authors
Recent
Search
2000 character limit reached

Learning Collapsed Patterns in Compositional Data: A Bayesian Heterogeneous Relative-Shift Approach

Published 5 Jun 2026 in stat.ME | (2606.07373v1)

Abstract: Relative-shift regression provides a principled framework for modeling compositional covariates by quantifying how the response changes when mass is reallocated from one component to another. Yet many emerging compositional data problems extend beyond this classical setting, involving high-dimensional predictors and regression effects that vary across latent subpopulations. This complexity poses a dual challenge unmet by existing methods: recovering latent cluster structure while simultaneously achieving dimension reduction within each cluster. We propose a Bayesian heterogeneous relative-shift regression model that jointly learns latent clusters and parsimonious effect structures. Methodologically, we combine a projection-based shrinkage prior on identifiable contrasts, which induces exact coefficient ties within mixture components, with a mixture of finite mixtures prior that infers the number of clusters. Computationally, we develop a scalable hybrid MCMC algorithm that embeds a deterministic surrogate collapse operator within NUTS. Theoretically, we establish posterior consistency for both the latent partition and cluster-specific effect structures. Simulations confirm accurate recovery and strong predictive performance, and applications to cross-country macroeconomic data and spatial transcriptomics demonstrate the method's interpretability and practical utility.

Authors (2)

Summary

  • The paper introduces a Bayesian heterogeneous relative‐shift regression model that fuses coefficient contrasts to reveal collapsed effect patterns.
  • It employs a projection‐based generalized ℓ1‐ball prior to stabilize inference and accurately recover latent clusters in high-dimensional settings.
  • Applications in macroeconomics and spatial transcriptomics demonstrate improved clustering, reduced MSE, and enhanced interpretability over standard models.

Learning Collapsed Patterns in Compositional Data: Bayesian Heterogeneous Relative-Shift Regression

Introduction

The paper "Learning Collapsed Patterns in Compositional Data: A Bayesian Heterogeneous Relative-Shift Approach" (2606.07373) develops a Bayesian framework for regression with high-dimensional compositional predictors, addressing the dual problems of latent subpopulation heterogeneity and the presence of collapsed (i.e., blockwise-constant) regression effects. By extending the relative-shift regression model to a mixture-of-regressions structure with direct regularization in the contrast space, and by leveraging projection-based priors, the method provides exact contrast-level aggregation and stabilizes high-dimensional posterior inference.

Methodological Framework

The foundation is the relative-shift regression model, which emphasizes identifiable regression contrasts rather than non-identifiable coefficient levels under simplex constraints. For observations yiy_i with compositional covariates xiSp1\mathbf{x}_i \in \mathcal{S}^{p-1}, the regression is parameterized so that xiθ\mathbf{x}_i^\top \boldsymbol{\theta} enters the mean function, but only pairwise differences θjθ\theta_j - \theta_\ell are identifiable. Direct interpretability comes from quantifying the response change under unit mass reallocation between components.

The model extends to the heterogeneous regime by introducing latent cluster indicators ziz_i for a mixture-of-regressions, so the composition--response relationship is allowed to vary across unobserved subpopulations. Cluster-specific coefficients θk\boldsymbol{\theta}_k inherit the simplex-induced nonidentifiability, preserving the need to regularize contrast space only. In high-dimensional settings, standard unconstrained or coordinatewise-shrinkage priors are shown to cause catastrophic posterior degeneracy: the partition collapses either to complete pooling or maximal separation, making accurate recovery of both clusters and collapsed structures impossible.

The proposed solution is to employ a projection-based generalized 1\ell_1-ball prior (GLBP), which deterministically projects unconstrained Gaussian precursors to the polyhedral set defined by upper bounds on the sum of contrast absolute values (i.e., total variation in the coefficients). By design, the polytope’s faces correspond to sets of tied coefficients, so the induced prior places positive probability exactly on collapsed (fused) patterns. Each mixture component’s aggregation structure is data-driven, with a separate collapse radius rkExp(ar)r_k \sim \mathrm{Exp}(a_r). The latent partition is modeled by a mixture-of-finite-mixtures (MFM) prior, which is preferred over the Dirichlet process due to its consistency in cluster number estimation.

Posterior Computation

The algorithmic underpinning is a hybrid MCMC scheme embedding structural regularization in the NUTS framework. Discrete cluster labels are marginalized for differentiability, and the projection (surrogate collapse operator) is implemented in closed form by soft-thresholding sorted coefficient differences. Label-switching is handled explicitly by co-clustering matrix alignment, using the Dahl partition as a point estimate with permutation-resolved averaging.

Theoretical Properties

The analysis establishes strong identifiability in the sense of second-order Fisher information separation, ensuring that both the mixing measure and cluster-specific effect structures are statistically learnable. Posterior consistency is proven: the posterior concentrates at optimal rates on both the true number of clusters and the true latent structure, overcoming the high-dimensional degeneracy of unregularized Bayesian mixtures.

Simulation Results

Empirical results on synthetic data demonstrate that the GLBP-based model outperforms baseline mixture-of-regressions with unconstrained or coordinatewise shrinkage priors. In all regimes, including those with pn/Kp \gtrsim n/K, the model shows:

  • High adjusted Rand indices (ARI) for partition recovery, often achieving perfect clustering.
  • Substantial reductions in coefficient mean squared error (MSE), especially as the ambient dimension pp increases.
  • Robustness under low sample sizes per cluster, preventing cluster fragmentation due to noise.

Application: Cross-Country Macroeconomic Analysis

The model is applied to 2019 macroeconomic data for 105 countries, regressing log PPP-adjusted GDP per capita on 18 ISIC sectoral shares plus non-compositional controls. The method discovers latent clusters with distinct sector-income relationships, but crucially, sector coefficients within clusters are collapsed to a small number of blocks, demonstrating blockwise-constant effect patterns.

Figure 1

Figure 1: Estimated cluster assignments for the 105 countries under the proposed model, based on the Dahl partition.

Notably, the incremental xiSp1\mathbf{x}_i \in \mathcal{S}^{p-1}0 from compositional predictors, after accounting for controls, is substantially lower with GLBP regularization compared to an unstructured baseline. This indicates that predictive power in sectors is captured largely by a few interpretable aggregated contrasts. Cluster sizes and average GDP differ only weakly, implying effect patterns are not driven exclusively by income.

Figure 2

Figure 2: Left: posterior co-fusion frequencies for sector pairs across clusters. Right: cluster-specific posterior mean coefficient values.

Within clusters, the effective model dimension is reduced from 18 to between 5 and 9 (Table 1 in the paper), showing strong parsimony without prediction loss. Co-fusion matrices reveal repeated and reproducible collapses of sector coefficients, reflecting interpretable structural aggregation in economic effects.

Application: Spatial Transcriptomics

The method is further validated on spatial transcriptomics from a human lung adenocarcinoma tissue section. For SPP1 gene expression, spot-level cell-type proportions are regressed with Poisson likelihood, revealing strong region-specific aggregation in cell-type effects and spatially contiguous clusters.

Figure 3

Figure 3: Estimated cluster assignments mapped onto the spatial coordinates of the lung adenocarcinoma tissue section. Each spot is colored by its Dahl-partition cluster; spatially contiguous patches of the smaller clusters indicate biologically and geographically structured heterogeneity in the composition--SPP1 relationship.

Figure 4

Figure 4: Fused cell type coefficients xiSp1\mathbf{x}_i \in \mathcal{S}^{p-1}1 across the SPP1 clusters, with groupings corresponding to identically fused coefficients.

Tumor cells and macrophages universally drive high SPP1 expression, but smaller clusters exhibit region-specific pattern shifts, e.g., ciliated epithelial cell effects emerging only locally. The fitted aggregation fuses large sets of weak or highly correlated cell types, reducing overfitting and enhancing biological interpretability.

Implications and Future Directions

This methodology provides a theoretically sound, computationally efficient mechanism for detection of latent clusters and interpretably collapsed effect structures in high-dimensional compositional regression, even when xiSp1\mathbf{x}_i \in \mathcal{S}^{p-1}2.

Notable implications include:

  • Stable inference under high-dimensionality: Direct regularization on contrasts, rather than on coefficients, is required for robust partition learning in mixture of regressions models.
  • Post hoc reliability: Exact coefficient fusion by projection—rather than soft penalties—aligns with interpretable effect identification and credible uncertainty quantification.
  • Applicability in genomics and macroeconomics: The model handles domain-specific features such as simplex constraints, high zero inflation, and known or hypothesized structural aggregation (e.g., sector hierarchies, taxonomies).

Potential future extensions involve elaborating the contrast operator to respect problem-specific hierarchies, integrating structured partitions (e.g., Markov random fields for spatial clustering), and accommodating interaction effects or more general dependence structures.

Conclusion

The Bayesian heterogeneous relative-shift regression framework with generalized xiSp1\mathbf{x}_i \in \mathcal{S}^{p-1}3-ball priors provides a principled solution for learning both latent heterogeneity and collapsed effect patterns in compositional data. The methodology delivers substantive improvements in clustering and parameter estimation over alternatives, supports interpretable effect grouping, and is theoretically justified to avoid degenerate inference regimes that plague high-dimensional unconstrained mixture models. This approach is broadly adaptable to compositional modeling tasks in economics, genomics, and any domain with simplex-structured predictors and unknown latent structure.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.