---
title: Compressed Bayesian Tensor Regression
url: https://www.emergentmind.com/topics/compressed-bayesian-tensor-regression
type: topic
---

# Compressed Bayesian Tensor Regression

Compressed Bayesian tensor regression denotes a family of Bayesian regression methods for tensor-valued data in which high-dimensional coefficient structures, covariates, or covariance components are replaced by lower-dimensional representations that preserve multiway structure while enabling posterior inference. In the cited literature, compression is achieved through low-rank CP/PARAFAC factorizations, Tucker decompositions, generalized tensor random projections, compressed random-effects covariance parameterizations, and sparse Kronecker product decompositions, depending on whether the setting is scalar-on-tensor, tensor-on-tensor, mixed-effects, or mixed-response regression [1509.06490]. Across these formulations, the common objective is to avoid naive vectorization of tensors, reduce the effective parameter dimension, and retain uncertainty quantification through Gibbs sampling or related MCMC schemes [1701.01037].

## 1. Scope and model classes

Compressed Bayesian tensor regression appears in several closely related regression settings. In scalar-on-tensor regression, a scalar response is modeled against a tensor predictor through an inner product with a coefficient tensor, as in
$$
y_i \mid \mathcal X_i,\boldsymbol{\beta},\sigma^2
\sim \mathcal{N}\!\Bigl(\langle \boldsymbol{\beta},\,\mathcal X_i\rangle,\;\sigma^2\Bigr),
$$
with \(\mathcal X_i\in\mathbb R^{p_1\times\cdots\times p_D}\) and \(\boldsymbol{\beta}\in\mathbb R^{p_1\times\cdots\times p_D}\) [1509.06490]. A related form retains the same scalar-on-tensor structure but first replaces the original tensor covariates by compressed tensors \(\widetilde{\mathcal X}_j\), yielding
$$
y_j = \mu + \bigl\langle\mathcal{B},\,\widetilde{\mathcal{X}_j}\bigr\rangle
+\sigma\,\varepsilon_j,
\quad \varepsilon_j\overset{\mathrm{iid}{\sim}}N(0,1),
$$
with \(m_1\cdots m_M\ll d_1\cdots d_K\) after projection [2510.01861].

In tensor-on-tensor regression, both predictor and response are tensors. One formulation assumes
$$
\mathbb Y \;=\;\bigl\langle\,\mathbb X\,,\,\mathbb B\,\bigr\rangle_{\!L}\;+\;\mathbb E,
$$
where \(\mathbb X\in\mathbb R^{N\times P_1\times\cdots\times P_L}\), \(\mathbb Y\in\mathbb R^{N\times Q_1\times\cdots\times Q_M}\), and \(\mathbb B\in\mathbb R^{P_1\times\cdots\times P_L\times Q_1\times\cdots\times Q_M}\) is the regression coefficient tensor [1701.01037]. This framework generalizes prediction of a scalar outcome from a tensor, a matrix from a matrix, or a tensor from a scalar [1701.01037].

Mixed-effects extensions incorporate repeated measures and random effects. CoMET assumes, for subject \(i\) and occasion \(j\),
$$
y_{ij} = \bigl\langle X_{ij}, \Bcal\bigr\rangle + \bigl\langle Z_{ij}, \Acal_i\bigr\rangle + \epsilon_{ij},
$$
with tensor-valued fixed-effect and random-effect covariates and tensor-normal random slopes [2602.19236]. A different mixed-effects construction treats the response itself as tensor-valued and introduces subject- and region-specific random effects to study both activation and connectivity in fMRI [1904.00148].

A further generalization handles mixed-type multivariate responses. In the BSKPD framework, each response component satisfies
$$
\theta_{ik} = \langle\mathcal X_i,\mathcal C_k\rangle + z_i^\top\gamma_k + u_{ik},
$$
with Gaussian, Poisson, and Bernoulli outcomes embedded in a common exponential-family form [2505.13821].

## 2. Compression mechanisms

A central distinction within the literature is where compression is applied. In one major line of work, compression is imposed directly on the regression coefficient tensor through low-rank structure. For scalar-on-tensor regression, the coefficient tensor is written in rank-\(R\) CP form,
$$
\boldsymbol{\beta}
=
\sum_{r=1}^R
u^{(1)}_r\circ u^{(2)}_r\circ\cdots\circ u^{(D)}_r,
$$
reducing the parameter count from \(\prod_{d=1}^D p_d\) to \(\sum_{d=1}^D R\,p_d\) [1509.06490]. In tensor-on-tensor regression, the same idea appears as a reduced CP-rank coefficient tensor
$$
\mathcal B
=
\sum_{r=1}^R
\bigl(a_{1r}\circ\cdots\circ a_{Lr}\circ b_{1r}\circ\cdots\circ b_{Mr}\bigr),
$$
with total free parameters
$$
R\bigl(\sum_{\ell=1}^L P_\ell+\sum_{m=1}^M Q_m\bigr),
$$
often orders of magnitude smaller than \(P\cdot Q\) [1701.01037].

A second line replaces CP by Tucker compression. In Bayesian tensor-on-tensor regression with efficient computation, the coefficient tensor is parameterized as
$$
\mathbb B
=
[\![\,\mathbb G\;;\;U_{(1)},\dots,U_{(L)},\,V_{(1)},\dots,V_{(M)}]\!],
$$
where the core tensor \(\mathbb G\) and factor matrices \(U_{(l)}\), \(V_{(m)}\) determine mode-specific multilinear ranks \((R_1,\dots,R_L,S_1,\dots,S_M)\) [2210.11363]. This permits modewise rank selection rather than a single CP rank.

A third mechanism compresses the predictors rather than only the coefficients. CBTR introduces a Generalized Tensor Random Projection in which a tensor \(\mathcal X\) is mapped to \(\widetilde{\mathcal X}=\mathrm{GTRP}(\mathcal X)\) through mode-wise projections, tensor-wise projections, or a combination of the two [2510.01861]. The paper explicitly identifies classic random projection on \(\mathrm{vec}(\mathcal X)\), separate mode-wise projection on each mode, and tensor-wise sketching that reduces the tensor order as special cases [2510.01861].

A fourth mechanism compresses covariance structure in mixed-effects models. In CoMET, each mode-\(d\) random-effects covariance factor is approximated using random matrices \(R_d,S_d\in\mathbb R^{k_d\times q_d}\) with \(k_d\ll q_d\), producing compressed random-effects covariates \(\tilde Z_{ij}\) and low-dimensional covariance parameters \(\Gamma_d\in\mathbb R^{k_d\times k_d}\) [2602.19236]. The resulting covariance parameter is low-dimensional enough to admit simple Gaussian priors [2602.19236].

A fifth form of compression uses a sparse Kronecker product decomposition. In BSKPD, each third-order coefficient tensor \(\mathcal C_k\) is represented through an unfolding
$$
\mathcal T(\mathcal C)=\sum_{r=1}^R \mathrm{vec}(A_r)\,\mathrm{vec}(B_r)^\top = BA^\top,
$$
which yields
$$
\langle\mathcal X_i,\mathcal C\rangle
=
\langle\mathcal T(\mathcal X_i),\mathcal T(\mathcal C)\rangle
=
\operatorname{tr}(A^\top \tilde X_i B)
=
\sum_{r=1}^R a_r^\top \tilde X_i b_r
$$
[2505.13821]. This preserves multilinear structure while decomposing the coefficient tensor into lower-dimensional factors.

These constructions indicate that “compression” in this area is not confined to one algebraic device. In the cited work, it can mean rank restriction, sketching, covariance projection, or structured factorization, provided that the resulting model remains Bayesian and tensor-aware.

## 3. Prior design and Bayesian regularization

Bayesian compression is coupled with prior distributions that stabilize estimation and often induce adaptive sparsity. In the CP-based scalar-on-tensor model, Guhaniyogi et al. introduce the M–DGDP prior. Component scales satisfy \(\tau_r=\phi_r\tau\), with \(\Phi=(\phi_1,\dots,\phi_R)\sim\mathrm{Dirichlet}(\alpha_1,\dots,\alpha_R)\) and \(\tau\sim\mathrm{Gamma}(a_\tau,b_\tau)\), while margin vectors satisfy
$$
u^{(d)}_r\mid \tau_r,W_{d,r}
\sim
\mathcal N(0,\tau_r W_{d,r}),
$$
with exponential local scales \(w_{d,r,k}\) and Gamma hyperpriors on \(\lambda_{d,r}\) [1509.06490]. Small \(\phi_r\) values effectively switch off redundant CP components [1509.06490].

In Lock’s tensor-on-tensor framework, the ridge-penalized least-squares criterion
$$
\widehat{\mathcal B}
=
\underset{\mathrm{rank}(\mathcal B)\le R}{\arg\min}\;
\|Y-\langle X,\mathcal B\rangle_L\|_F^2
+
\lambda\|\mathcal B\|_F^2
$$
is shown to be the mode of a Bayesian posterior under a Gaussian likelihood and a Gaussian prior on \(\mathrm{vec}(\mathcal B)\) supported on rank \(\le R\) [1701.01037]. Here the compression is low rank and the regularization is ridge-like, rather than sparse in the usual global-local sense.

The Tucker-based BayTensor model employs conjugate Gaussian-inverse-gamma priors on the core tensor, factor matrices, and noise variance:
\(\operatorname{vec}(\mathbb G)\sim N(\mu_G,\Sigma_G)\),
\(\operatorname{vec}(U_{(l)})\sim N(\mu_U,\Sigma_U)\),
\(\operatorname{vec}(V_{(m)})\sim N(\mu_V,\Sigma_V)\),
and \(\sigma^2\sim\mathrm{IG}(\alpha,\beta)\), with diagonal covariance matrices and a prior \(\pi(\theta)\) on the model dimensions \(\theta=(R_1,\dots,R_L,S_1,\dots,S_M)\) [2210.11363]. This separates parameter shrinkage from model-dimension uncertainty.

CoMET combines low-rank CP structure for fixed effects with Horseshoe priors. Each entry \(\beta^{(g)}_{dj}\) receives
$$
\beta^{(g)}_{dj}\mid\lambda_{gdj},\delta_g,\tau
\sim\mathcal N(0,\lambda_{gdj}^2\delta_g^2\tau^2),
$$
with local half-Cauchy scales \(\lambda_{gdj}\), component-level half-Cauchy scales \(\delta_g\), and inverse-gamma prior on \(\tau^2\) [2602.19236]. The paper states that this supports both parameter reduction and sparsity and enables fixed-effects selection [2602.19236].

Spencer et al. use a multiway stick-breaking shrinkage prior for tensor-response regression. With
$$
\phi_{g,r}=\xi_{g,r}\prod_{l<r}(1-\xi_{g,l}),
\qquad
\xi_{g,r}\sim\mathrm{Beta}(1,\alpha_g),
$$
the component weights satisfy an ordered shrinkage pattern \(\phi_{g,1}\ge\phi_{g,2}\ge\cdots\), while local exponential mixtures induce generalized double-Pareto-type voxelwise shrinkage [1904.00148].

BSKPD adopts Three-Parameter Beta-Normal shrinkage priors for coefficient-factor matrices and auxiliary covariates, together with an inverse-Wishart prior for the residual covariance \(\Sigma\) [2505.13821]. This construction is explicitly designed for voxel-level sparsity and interpretability in mixed-type response models [2505.13821].

CBTR, after projection, returns to a PARAFAC coefficient representation and a hierarchical global-local shrinkage prior inherited from Guhaniyogi et al. The margin vectors \(\boldsymbol{\gamma}_m^{(d)}\) follow Gaussian priors with component-wise and global scales, local exponential scales, Gamma hyperpriors, inverse-gamma prior on \(\tau\), and Dirichlet prior on the component weights \(\boldsymbol{\zeta}\) [2510.01861].

## 4. Posterior computation and optimization

The computational core of compressed Bayesian tensor regression is typically a block Gibbs sampler exploiting conjugate or conditionally Gaussian updates. In the rank-restricted tensor-on-tensor model, under the CP parameterization \(\mathcal B=\llbracket U_1,\dots,U_L,V_1,\dots,V_M\rrbracket\), each block \(U_\ell\) or \(V_m\) has a conditionally Gaussian full conditional. For example,
$$
\mathrm{vec}(U_1)\mid \text{rest}
\sim
N(m_1,\Sigma_1),
$$
with
$$
m_1=
\bigl(C^\top C+\lambda G_1\otimes I_{P_1}\bigr)^{-1}C^\top\mathrm{vec}(Y),\qquad
\Sigma_1=
\sigma^2\bigl(C^\top C+\lambda G_1\otimes I_{P_1}\bigr)^{-1},
$$
and \(\sigma^2\) updated from an inverse-gamma conditional [1701.01037]. The same paper also describes alternating least squares on the penalized objective, with warm starts at large \(\lambda\) followed by decreasing \(\lambda\), described there as “tempered regularization” [1701.01037].

In BayTensor, posterior inference alternates between Gibbs updates conditional on fixed multilinear rank \(\theta\) and a rank–Metropolis–Hastings step on neighboring rank configurations satisfying \(\|\tilde\theta-\theta\|_1=1\) [2210.11363]. Rather than full Reversible Jump MCMC, the method uses a fractional-Bayes-factor Metropolis–Hastings move to jointly update dimension and associated parameters [2210.11363]. The same work also provides an optimization-based alternative: closed-form MAP updates for \(\{U_{(l)},V_{(m)},G,\sigma^2\}\) at fixed \(\theta\), combined with simulated annealing over \(\theta\) using a BIC objective [2210.11363].

The M–DGDP model likewise uses a blocked Gibbs sampler. The main blocks are the rank weights \((\Phi,\tau)\), the margin-specific parameters \(\{u^{(d)}_r,w_{d,r,k},\lambda_{d,r}\}\), and \(\sigma^2\), with generalized inverse-Gaussian, Gamma, multivariate normal, and inverse-gamma updates arising from conjugacy and scale-mixture representations [1509.06490].

CoMET uses a two-block collapsed Gibbs sampler. One block samples compressed random effects \(\tilde d_i\) and covariance parameters \(\gamma_d=\mathrm{vec}(\Gamma_d)\); the other marginalizes over \(\tilde d_i\), whitens the resulting regression blocks, and samples the mode-specific factor margins \(\tilde\beta_d=\mathrm{vec}(B_d)\), local shrinkages, global shrinkages, and \(\tau^2\) from closed-form conditionals [2602.19236]. The paper states that all full-conditionals are in closed form and that each sweep costs only \(O(NK\sum p_d + Nk^2)\), roughly linear in the original tensor dimensions [2602.19236].

For mixed-type outcomes, BSKPD combines TPBN shrinkage with Polya-Gamma augmentation. Binary responses are converted to a Gaussian-form conditional likelihood using \(\omega\sim PG(1,0)\) and \(\kappa=y-1/2\), after which \(A_k\), \(B_k\), \(\gamma_k\), \(U\), and \(\Sigma\) admit closed-form Gibbs updates [2505.13821].

CBTR also uses a custom Gibbs sampler with Gaussian updates for the margin vectors, generalized inverse-Gaussian and inverse-gamma updates for hyperparameters, and Gaussian/inverse-gamma updates for \(\mu\) and \(\sigma^2\) [2510.01861]. To reduce sensitivity to a single projection, it further employs Bayesian model averaging over multiple independent random projections, with model evidences estimated by reverse logistic regression [2510.01861].

## 5. Theory: concentration, consistency, and dimension recovery

Several strands of theory support compressed Bayesian tensor regression, though the precise guarantees depend on the compression mechanism. In the scalar-on-tensor CP model with M–DGDP prior, posterior consistency holds under mild regularity, including a true coefficient tensor of rank-\(R\) CP form, bounded true margins, and growth condition \(\sum_{d=1}^D p_{d,n}\log p_{d,n}=o(n)\) [1509.06490]. The paper states that the posterior concentrates on any Kullback–Leibler neighborhood of the true \(\boldsymbol{\beta}\) at the usual root-\(n\) rate up to logarithmic factors [1509.06490].

CBTR adds theory for random projection itself. Using projection entries drawn i.i.d. from the “database-friendly” distribution of Achlioptas, the paper proves Johnson–Lindenstrauss-type distance preservation: provided the projected dimension satisfies \(q=m_1\cdots m_M\gtrsim \epsilon^{-2}\log n\), pairwise Frobenius distances among tensors are preserved within factors \(1\pm\epsilon\) with high probability [2510.01861]. Under additional regularity conditions, Theorem 4.1 and Theorem 4.2 establish posterior contraction for Gaussian and PARAFAC priors after compression [2510.01861].

For Tucker tensor-on-tensor regression, the principal theoretical contribution is inferential rather than asymptotic: the MCMC sampler jointly estimates both model dimension and parameters, addressing a limitation of earlier Tucker methods that either assumed the core dimension to be known or selected it through cross-validation or model selection criteria [2210.11363]. The empirical section further reports Bayesian dimension recovery rates of 80–100% for MCMC and 40–80% for the simulated-annealing alternative in the simulated settings described in the paper [2210.11363].

CoMET provides a high-dimensional predictive guarantee. Under assumptions labeled (A1)–(A3), including bounded operator norms, sub-Gaussian covariates, and compression-rank conditions such as \(k^{*2}\log k^*\ll p^*\) and \(p^*\ll N\), the posterior predictive mean \(\bar\beta\) satisfies
$$
\frac1N\EE\bigl\|X\beta^0-X\bar\beta\bigr\|_2^2=o(1),
$$
so the posterior predictive risk tends to zero [2602.19236].

BSKPD proves both identifiability and posterior consistency. Theorem 1 establishes asymptotic identifiability of the sparse Kronecker product decomposition up to the usual orthogonal non-identifiability in \(A,B\), and Theorem 2 shows posterior concentration on Frobenius-norm neighborhoods of \((C_0,\gamma_0)\) in both low-dimensional and high-dimensional regimes under stated sparsity and growth conditions [2505.13821].

Taken together, these results suggest that compression in Bayesian tensor regression is not treated solely as a computational heuristic. In these papers it is paired with explicit guarantees on posterior concentration, dimension recovery, distance preservation, or predictive risk.

## 6. Empirical behavior, applications, and limitations

The empirical literature emphasizes prediction, uncertainty quantification, and parameter savings. In Lock’s facial-attribute experiment, images were represented as \(90\times90\times3\) RGB tensors and outcomes as 72 standardized attributes, with \(N=1000\) training faces and \(N=1000\) held-out faces [1701.01037]. A grid search over \(R=1,\dots,16\) and \(\lambda\) up to \(10^5\) selected \(R=15\), \(\lambda=10^5\), yielding relative prediction error approximately \(0.568\), compared with greater than \(0.8\) for unregularized or full-rank ridge, and overall correlation approximately \(r=0.66\) between observed and predicted attributes [1701.01037]. Posterior sampling with \(T=5000\) draws gave out-of-sample 95% credible-interval coverage approximately \(0.93\) and 90% coverage approximately \(0.887\) [1701.01037].

The Tucker-based BayTensor method reports stronger performance on related tasks. On facial-image-to-attribute data with \(1000\) images and \(90\times90\times3\to73\) labels, BayTensor MCMC uses approximately \(1\,154\) parameters and attains RPE \(=0.38\), whereas the CP comparator uses \(3\,840\) parameters with RPE \(=0.48\) [2210.11363]. On MoCap data with \(200\) frames and \(32\times24\to3\times37\), BayTensor MCMC has RPE approximately \(0.18\) versus \(0.25\) for CP [2210.11363]. Posterior predictive credible intervals achieve approximately \(94\)–\(96\%\) empirical coverage for nominal \(95\%\) intervals in simulated and real-data settings [2210.11363].

In the mixed-effects tensor-response fMRI model of Spencer et al., rank \(3\) minimized DIC, achieved root-MSE around \(0.08\) versus \(0.11\) for the vectorized model, yielded activation detection AUC approximately \(0.95\) versus approximately \(0.91\), and used \(927\) parameters versus \(11\,078\) in the vectorized model [1904.00148]. In the Balloon-Analog Risk-Taking application, the rank-\(3\) model also had healthy MCMC ESS \(>800\) and produced posterior partial-correlation patterns consistent with known risk-processing circuits [1904.00148].

CBTR’s simulation study varies tensor size, sample size, projection scheme, sparsity of the projection, and PARAFAC rank. The reported findings are that mode-wise GTRP outperforms tensor-wise in most settings, moderate sparsity \(\psi\approx3\) is often optimal, and model averaging over \(L=10\) projections stabilizes predictions [2510.01861]. In the oil-volatility and S&P 500 application, CBTR-MW and CBTR-MW(1,2) outperform uncompressed BTR out of sample with RMSE \(0.038\) versus \(0.115\), while maintaining similar in-sample fits [2510.01861].

CoMET is reported to outperform penalized competitors across simulation studies and benchmark applications involving facial-expression prediction and music emotion modeling [2602.19236]. The paper attributes tractability to structured, mode-wise random projection of the random-effects covariance and a collapsed Gibbs sampler whose computational complexity grows approximately linearly with the tensor covariate dimensions [2602.19236].

The literature also identifies limitations. BayTensor notes that very large tensors may require more advanced samplers such as variational or pseudo-marginal methods [2210.11363]. CBTR observes sensitivity to the choice of random projection and addresses this by Bayesian model averaging; it also notes that loss of interpretability may occur if projection is too aggressive [2510.01861]. These points counter a common misconception that compression is uniformly benign: the cited work instead treats compression as a tunable statistical approximation whose benefits depend on preserving relevant tensor structure.

## 7. Relationship to adjacent methods

Compressed Bayesian tensor regression is closely related to reduced-rank regression, tensor decomposition, structured shrinkage, and high-dimensional Bayesian regression, but it is distinguished by its simultaneous use of tensor algebra and posterior uncertainty quantification. Lock’s tensor-on-tensor regression explicitly unifies tensor-on-tensor, reduced-rank, and ridge-penalized models in a single framework [1701.01037]. Guhaniyogi et al. connect tensor regression to parafac decomposition, multiway shrinkage, and posterior consistency in growing dimensions [1509.06490]. Wang and Xu extend the tensor-on-tensor setting by replacing CP with Tucker decomposition and by estimating model dimension inside the Bayesian algorithm rather than by external cross-validation [2210.11363].

Random-projection approaches such as CBTR differ from low-rank coefficient models in that they compress the covariates before or during inference, with explicit Johnson–Lindenstrauss-type guarantees and Bayesian model averaging across sketches [2510.01861]. Mixed-effects formulations such as CoMET show that compression can be applied not only to coefficient tensors but also to random-effects covariance structure [2602.19236]. BSKPD indicates that a low-rank Kronecker product decomposition can support unified Gaussian, Poisson, and Bernoulli responses while maintaining closed-form Gibbs updates through Polya-Gamma augmentation [2505.13821].

A plausible implication is that the field has moved from a relatively narrow focus on scalar-on-tensor CP models toward a broader class of compressed Bayesian tensor regressions in which the point of compression is chosen to match the inferential bottleneck: coefficient dimensionality, unknown multilinear rank, predictor size, random-effects covariance, or mixed-response coupling. The common design principle remains the same: preserve multiway structure, reduce effective dimension, and maintain a fully probabilistic inferential pipeline.

Source: https://www.emergentmind.com/topics/compressed-bayesian-tensor-regression