---
title: Location-Dependent Stick-Breaking Prior
url: https://www.emergentmind.com/topics/location-dependent-stick-breaking-prior
type: topic
---

# Location-Dependent Stick-Breaking Prior

A location-dependent stick-breaking prior is a family of random probability measures indexed by a predictor, location, or space-time coordinate, typically written \(G_x=\sum_h w_h(x)\delta_{\theta_h}\), in which the atoms are shared across index values while the stick-breaking weights vary with \(x\). In this usage, “location” may mean a baseline covariate vector \( \bm z \), a spatial coordinate \(s\), a spatio-temporal index \((s,t)\), or a more general input \(z\); the defining feature is that the mixing distribution changes with the index through \(w_h(x)\), not that the likelihood alone depends on covariates [2203.12280; 2303.17177; 1701.02969]. This places location-dependent stick-breaking within the broader class of dependent random probability measures, and, in the common-atoms form, it yields predictor- or location-dependent clustering through index-specific mixture probabilities rather than through index-specific atoms alone [2208.02806; 1506.05843].

## 1. Definition and formal structure

In the ordinary Sethuraman representation of a Dirichlet process,
\[
F(\cdot)=\sum_{k=1}^{\infty}\pi_k\delta_{\theta_k}(\cdot), \qquad \pi_1=V_1,\quad \pi_k=V_k\prod_{j=1}^{k-1}(1-V_j),\ k\ge2,
\]
the weights do not depend on any covariate or index value. Every observation sees the same \(\{\pi_k\}\). A location-dependent stick-breaking prior replaces this exchangeable construction by an indexed family
\[
G_x=\sum_{h} w_h(x)\delta_{\theta_h},
\qquad
w_h(x)=V_h(x)\prod_{\ell<h}\{1-V_\ell(x)\},
\]
so that the random measure changes with \(x\) through the break proportions or their transforms [2303.17177; 1701.02969].

The common-atoms form is especially prominent. In the child-growth VAR model,
\[
\Phi_i \mid z_i \sim \sum_{h=1}^H w_h(\bm z_i)\delta_{\Phi_{0h}},
\]
with global atoms \(\Phi_{0h}\) shared across subjects and covariate-dependent weights \(w_h(\bm z_i)\) [2203.12280]. In the spatio-temporal formulation,
\[
F_{s,t}(\cdot)=\sum_{k=1}^{\infty}\pi_k(s,t)\,\delta_{\theta_k}(\cdot),
\]
the same countable atoms are used for all \((s,t)\), while the weight vector varies with location and time [2303.17177]. The same principle appears in covariate-dependent treeSB models,
\[
G_x = \sum_{\epsilon\in B(\tau)} W_{x,\epsilon}\,\delta_{\theta_\epsilon},
\]
where each leaf weight is a path probability determined by covariate-dependent binary gates [2208.02806].

This construction is distinct from models in which covariates affect only the kernel or only the atom locations. The sources repeatedly separate three mechanisms: common weights with varying atoms, varying weights with common atoms, and hybrids in which both vary. In location-dependent stick-breaking priors, the defining dependence is in the weights. This is why they are often described as predictor-dependent mixtures, spatial stick-breaking priors, or spatio-temporal stick-breaking processes, depending on the index set [2203.12280; 2303.17177].

A useful latent-allocation view introduces \(c_i\) or \(G_i\) such that
\[
c_i\mid x_i \sim \mathrm{Categorical}(\{1,\ldots,H\};\bm w(x_i)),
\]
followed by a component-specific draw. This makes explicit that the prior acts on the random partition by changing allocation probabilities as the location index changes [2203.12280; 1701.02969].

## 2. Principal constructions

Several non-equivalent constructions instantiate the same general idea: keep the stick-breaking normalization but let the breaks depend on predictors, locations, or paths in a tree.

| Construction | Index | Dependence mechanism |
|---|---|---|
| Logit stick-breaking prior | \(\mathbf x\), \(\mathbf z\) | \(\operatorname{logit}\{\nu_h(x)\}=\psi(x)^\top\alpha_h\) |
| Spatial or spatio-temporal stick-breaking | \(s\), \((s,t)\) | \(V_k(s,t)=w_k(s,\psi_k,t,\zeta_k)V_k\) |
| Tree stick-breaking | \(x\) | Node-specific binary regressions along tree paths |
| Finite multinomial logistic stick-breaking | \(z\) | \(\pi_k(z)=\sigma(\psi_k(z))\prod_{j<k}(1-\sigma(\psi_j(z)))\) |

The logit stick-breaking prior is the most direct regression-based version. In the LSBP density-regression model,
\[
\pi_h(\mathbf x)=\nu_h(\mathbf x)\prod_{l=1}^{h-1}\{1-\nu_l(\mathbf x)\},
\qquad
\eta_h(\mathbf x)=\operatorname{logit}\{\nu_h(\mathbf x)\}=\psi(\mathbf x)^\top \alpha_h,
\]
with \(\alpha_h\sim N_R(\mu_\alpha,\Sigma_\alpha)\) [1701.02969]. The same mechanism appears in finite form in the obesity application,
\[
\text{logit}(\nu_h(\bm z_i))=\bm z_i^\top\bm\alpha_h,\qquad h=1,\dots,H-1,\qquad \nu_H(\bm z_i)=1,
\]
which yields a covariate-dependent finite stick-breaking prior for subject-specific VAR coefficients [2203.12280].

Spatial and spatio-temporal models instead modulate the breaks through kernels. In the spatial case,
\[
V_k(s)=w_k(s,\psi_k)V_k,
\]
and in the spatio-temporal extension,
\[
V_k(s,t)=w_k(s,\psi_k,t,\zeta_k)V_k,
\]
with \(V_k\sim\operatorname{Beta}(a,b)\), \(\psi_k\sim Q\), and \(\zeta_k\sim L\) [2303.17177]. The kernel may be separable,
\[
w_k(s,\psi_k,t,\zeta_k)=w_k(s,\psi_k)\cdot w_k(t,\zeta_k),
\]
or non-separable, as in the Gneiting-type example
\[
w(s,\psi,t,\zeta) = \frac{1}{\gamma |t-\zeta| + 1} \exp\!\left( -\frac{(s_1-\psi_1)^2+(s_2-\psi_2)^2}{(\gamma |t-\zeta| + 1)^{\lambda/2}} \right),
\]
with \(\lambda=0\) corresponding to separability [2303.17177].

Tree-based generalizations reinterpret stick-breaking as a bifurcating routing architecture. For a binary string \(\epsilon_1\cdots\epsilon_m\),
\[
W_{\epsilon_1 \cdots \epsilon_m}
=
\prod_{l=1}^m
V_{\epsilon_1\cdots \epsilon_{l-1}}^{\,1-\epsilon_l}
\bigl(1-V_{\epsilon_1\cdots \epsilon_{l-1}}\bigr)^{\epsilon_l},
\]
and, in the covariate-dependent form,
\[
V_{x,\epsilon} = \text{logistic}(\eta_{x,\epsilon}),
\qquad
\eta_{x,\epsilon} = \psi(x)^\top \gamma_{\epsilon},
\qquad
\gamma_{\epsilon} \sim N_R(\mu_{\epsilon}, \Sigma_{\epsilon})
\]
[2208.02806]. This formulation subsumes the standard sequential construction as a lopsided tree and introduces balanced trees as alternative topologies.

A related but finite-\(K\) construction appears in dependent multinomial models:
\[
\pi_k(z)=\sigma(\psi_k(z))\prod_{j<k}\bigl(1-\sigma(\psi_j(z))\bigr),
\qquad
\pi_K(z)=\prod_{j=1}^{K-1}\bigl(1-\sigma(\psi_j(z))\bigr),
\]
with latent Gaussian structure on \(\psi_k(z)\) over documents, time, space, or covariates [1506.05843]. Although this is not a BNP prior over infinitely many atoms, it is a genuine location-dependent stick-breaking model on the simplex.

## 3. How location dependence changes clustering and dependence structure

The practical effect of location dependence is that allocation probabilities vary systematically across the index. In the longitudinal VAR model,
\[
P(c_i=h\mid \bm z_i)=w_h(\bm z_i),
\]
so two children with different baseline covariates have different prior probabilities of belonging to the same latent autoregressive regime [2203.12280]. The paper emphasizes that the prior is specified on the random partition of patients, which implies that posterior clusters can be driven by responses, covariates, or both; accordingly, “number of clusters” need not equal “number of distinct trajectory profiles” [2203.12280].

In spatial and spatio-temporal stick-breaking, dependence is localized through kernel overlap. For the single-atom model,
\[
\mathbb{C}\mathrm{ov}\bigl(y(s,t),y(s',t')\mid V,\psi,\zeta\bigr)
=
\Pr\!\bigl(\theta(s,t)=\theta(s',t')\mid V,\psi,\zeta\bigr)
=
\sum_{k=1}^{\infty}\pi_k(s,t)\pi_k(s',t'),
\]
so nearby \((s,t)\) and \((s',t')\) have larger coincidence probability because the local kernels overlap more strongly [2303.17177]. This yields smooth index dependence without requiring Gaussian-process atoms in the baseline single-atom DDP construction. The same source notes, however, that single-atom DDPs have a positive lower bound on dependence; this motivates an extension with both \(\pi_k(s,t)\) and \(\theta_k(s,t)\) varying [2303.17177].

A recurring conceptual distinction concerns whether dependence enters through weights or atoms. Linear-DDP-type competitors may let atoms depend on covariates while leaving weights common, whereas logit stick-breaking priors change the weights directly [2203.12280]. Spatial DDPs may instead keep the weights common and vary \(\theta_k(s)\). The common-atoms, varying-weights form is canonical precisely because it changes local partition probabilities while preserving a shared global dictionary [2303.17177; 2208.02806].

Tree topology further alters the prior dependence structure. The lopsided tree underlying ordinary sequential stick-breaking induces strong cross-covariate prior correlation and inherited stochastic ordering, whereas the balanced tree produces a symmetric, shallower path representation [2208.02806]. For a balanced tree with \(K=2^m\),
\[
a_{x,x'} = \left\{ 1-E(V_x)-E(V_{x'})+2E(V_xV_{x'}) \right\}^m,
\]
so cross-location dependence can decay with depth, while local continuity is preserved if
\[
E(V_{x'}) \to E(V_x)
\quad\text{and}\quad
E(V_xV_{x'}) \to E(V_x^2)
\quad\text{as }x'\to x
\]
[2208.02806]. This directly links topology to the degree of prior smoothing across locations.

A related misconception concerns the meaning of “location.” In some papers it is spatial position; in others it is a predictor vector. The obesity application states this explicitly: the “location” index is the subject’s baseline covariate vector \( \bm z_i \), not physical space [2203.12280]. The general notion is therefore index dependence in the mixing weights, not necessarily geostatistical locality.

## 4. Posterior computation and algorithmic strategies

The computational appeal of location-dependent stick-breaking priors lies in the fact that several important constructions reduce the dependent-weight problem to conditionally Gaussian updates.

The central device is Pólya–Gamma augmentation. In the LSBP model, the sequential logistic representation of the breaks yields binary regression likelihoods for the continuation-ratio probabilities, and introducing \(\omega_{ih}\sim \mathrm{PG}(1,\psi(\mathbf x_i)^\top\alpha_h)\) makes the full conditional for \(\alpha_h\) Gaussian [1701.02969]. The same idea is used in the finite VAR clustering model, where the full conditional for the weight parameters \(\{\bm\alpha_h\}\) is derived in closed form using auxiliary variables following Polson, Scott and Windle (2013) and Rigon and Durante (2021), again via Pólya–Gamma augmentation [2203.12280]. In finite multinomial logistic stick-breaking, the multinomial decomposes into \(K-1\) binomial logistic terms,
\[
\mathrm{Multinomial}(\bx\mid N,\bpsi)
=
\prod_{k=1}^{K-1}
{N_k \choose x_k}
\frac{(e^{\psi_k})^{x_k}}{(1+e^{\psi_k})^{N_k}},
\]
so conditioned on \(\omega_k\sim \mathrm{PG}(N_k,\psi_k)\), the likelihood becomes Gaussian in \(\bpsi\) [1506.05843].

Truncation is another standard tactic. The obesity model uses finite truncation with \(H=25\) in simulations and \(H=50\) in the application [2203.12280]. LSBP likewise imposes \(\nu_H(\mathbf x)=1\) in the truncated approximation [1701.02969]. The spatio-temporal stick-breaking process uses a conditional MCMC approximation with \(M=100\) in simulations and applications [2303.17177]. These are blocked or conditional truncation schemes in the style of Ishwaran–James and related conditional samplers.

The latent-allocation step always couples local weights to local likelihood contributions. In the VAR model,
\[
P(G_i=h \mid \text{rest}) \propto w_h(\bm z_i)\, p(\bm y_i\mid \Phi_{0h},B,\Gamma,\Sigma),
\]
so covariate dependence affects clustering exactly through the weight factor [2203.12280]. In the LSBP Gaussian density-regression model,
\[
\Pr(G_i=h\mid -) \propto
\left[\nu_h(\mathbf x_i)\prod_{l=1}^{h-1}\{1-\nu_l(\mathbf x_i)\}\right]
\sqrt{\tau_h}\phi[\sqrt{\tau_h}\{y_i-\lambda(\mathbf x_i)^\top\beta_h\}],
\]
and analogous local-allocation probabilities appear in spatial and tree-based models [1701.02969; 2303.17177; 2208.02806].

Several constructions admit more than one inference regime. LSBP provides Gibbs sampling, expectation-maximization, and mean-field variational Bayes, all based on the same Pólya–Gamma augmentation [1701.02969]. The multinomial logistic stick-breaking model combines the augmentation with Gaussian-process regression, LDS smoothing, and block Gaussian updates [1506.05843]. TreeSB models preserve the same node-wise binary-regression decomposition under alternative tree topologies, so the move from lopsided to balanced trees does not destroy conditional-conjugate regression machinery [2208.02806].

Other algorithmic strategies are more specialized. The spatio-temporal stick-breaking process uses latent Bernoulli variables
\[
A_{ik}\sim \mathrm{Bern}(V_k),\qquad B_{ik}\sim \mathrm{Bern}(w(s,t,\psi,\zeta)),
\qquad H_i=\min\{k: A_{ik}=B_{ik}=1\},
\]
which separates the global beta variable \(V_k\) from the local kernel activation probability \(w_k\); knot variables \(\psi_k\) and \(\zeta_k\) are then updated with Metropolis–Hastings [2303.17177]. This illustrates a broader pattern: local dependence in the weights often requires auxiliary-variable designs that decouple global break magnitudes from local gating.

## 5. Theoretical properties and adjacent models

A basic requirement is normalization. For LSBP, the weights satisfy
\[
\sum_{h=1}^\infty \pi_h(\mathbf x)=1 \quad \text{a.s. for every } \mathbf x,
\]
and the paper proves this under the condition that \(\mu_\nu(\mathbf x)=E\{\nu_h(\mathbf x)\}\in(0,1)\) [1701.02969]. The same paper also gives an \(L^1\) truncation bound:
\[
\|f_{\mathbf X}^{(H)}(\mathbf y)-f_{\mathbf X}^{(\infty)}(\mathbf y)\|_1 \le 4\sum_{i=1}^n \{1-\mu_\nu(\mathbf x_i)\}^{H-1},
\]
so the finite approximation error decays exponentially in \(H\) [1701.02969].

For the spatio-temporal stick-breaking process, the weights are proper in the sense that
\[
\sum_{k=1}^{\infty}\pi_k(s,t)=1 \qquad \text{a.s.}
\]
provided \(\mathbb E[V_k]\) and \(\mathbb E[w_k(s,\psi_k,t,\zeta_k)]\) are positive [2303.17177]. Under these mild conditions, invoking Barrientos, Jara and Quintana (2012), the random measures are marginally DP-distributed for each \((s,t)\), the process has full weak support, and the induced DP mixture model has smooth trajectories as \((s,t)\) varies [2303.17177].

Covariance structure is more subtle than the kernel form alone might suggest. A separable weight kernel,
\[
w_k(s,\psi_k,t,\zeta_k)=w_k(s,\psi_k)\cdot w_k(t,\zeta_k),
\]
does not imply a separable covariance for the induced process because the stick-breaking product structure is nonlinear [2303.17177]. This is one of the clearest examples of how local dependence in weights produces nontrivial second-order behavior.

A broader theoretical backdrop comes from work that is not itself location-dependent. Exchangeable-length-variable stick-breaking processes generalize classical independent-stick priors and give properness and full-support criteria for dependent lengths, but they use one global weight sequence rather than \(x\)-indexed weights [2008.04475]. Markov stick-breaking processes replace independent lengths by a Markov chain and show that dependence among \(V_h\) can be tuned while preserving prescribed marginals, again without introducing a location index [2601.16561]. These models are relevant background because they separate marginal stick laws from dependence mechanisms, a design principle that is directly transferable to location-dependent priors.

Two nearby constructions clarify what does not count, or counts only indirectly, as location-dependent stick-breaking. The multiscale Bernstein polynomial prior is a tree-structured stick-breaking density model with localized beta basis functions, but its weights are indexed by tree node \((s,h)\), not by an external predictor or spatial coordinate [1410.0827]. The \(\psi\)-stick-breaking model for related samples couples mixture weights across discrete groups \(j\) through shared and idiosyncratic sticks, but it is group-dependent rather than spatially or covariate-indexed, and its “location sensitivity” enters mainly through kernel perturbation of the atoms [1704.04839]. These contrasts help delimit the subject: location-dependent stick-breaking priors are specifically those in which the stick variables or the resulting weights vary with an external index.

## 6. Applications, advantages, and limitations

The applied literature shows that location-dependent weights are useful when cluster proportions or local mixture composition are expected to vary systematically with baseline factors, spatial position, or time. In the child-obesity application, the prior clusters children through subject-specific VAR coefficient matrices \(\Phi_i\), after adjusting for baseline fixed effects and a global growth trend; the authors report better predictive performance relative to a purely parametric model (\(H=1\)), a DP prior on \(\Phi_i\) without covariate dependence, and a model where covariates affect atoms linearly but not weights, and they state that the reported WAIC values favored their logit stick-breaking model [2203.12280].

In density regression, the LSBP toxicology application models the conditional distribution of gestational age at delivery given maternal serum DDE concentration using \(n=2312\) observations, a Gaussian mixture kernel with
\[
\lambda_1(\bar x_i)=1,\qquad \lambda_2(\bar x_i)=\bar x_i,
\]
and a natural cubic spline basis in the stick-breaking logits with truncation \(H=20\) [1701.02969]. The fitted conditional densities show increasing left-tail inflation with higher DDE, and Gibbs, EM, and VB give similar substantive results; EM is fastest for point estimation, VB is much faster than MCMC while retaining uncertainty quantification, and Gibbs gives exact posterior inference at greater computational cost [1701.02969].

In spatial and spatio-temporal prediction, the empirical gains can be large. In one simulation, the benchmark separable Gaussian spatio-temporal model had ESPE \(175.69\), the spatial stick-breaking competitor with temporally evolving atoms had \(196.04\), and the proposed single-atom spatio-temporal stick-breaking had \(52.33\) [2303.17177]. On Australian rainfall data, using observations from 2007–2016 to predict 2017, the ESPEs were \(109.74\) for the Gaussian spatio-temporal model, \(112.09\) for spatial stick-breaking, and \(71.05\) for single-atom stSB; on California temperature data, the ESPEs were \(103.66\), \(123.67\), and \(40.53\), respectively [2303.17177]. The paper reports that spatial-only models oversmoothed, whereas stSB better captured local peaks, changing regional patterns, and recent local changes [2303.17177].

Finite location-dependent stick-breaking on the simplex is likewise effective in structured multinomial settings. The multinomial GP and multinomial LDS models use latent Gaussian priors on stick logits \(\psi_k(z)\) over year, latitude, longitude, or latent state trajectories, and the paper reports that the GP multinomial model is comparable in prediction to logistic-normal GP but considerably more efficient computationally, while the multinomial LDS is orders of magnitude faster than logistic-normal LDS with particle MCMC [1506.05843].

Several advantages recur across these applications. Predictor-dependent clustering allows systematic shifts in prior cluster membership as covariates change [2203.12280]. Kernel-based localization allows local borrowing of strength while avoiding the oversmoothing of kriging-like Gaussian processes and the rigidity of purely spatial stick-breaking models that cannot adapt in time [2303.17177]. Binary-regression formulations preserve computational tractability through Pólya–Gamma augmentation [1701.02969; 2208.02806].

The limitations are equally consistent. Truncation levels such as \(H\) or \(M\) must be chosen [2203.12280; 2303.17177]. Prior sensitivity can be substantial, especially for \(\Sigma_\alpha\) and atom-variance hyperparameters; vague priors can lead to extreme imputations and poor mixing [2203.12280]. In weight-dependent clustering models, clusters may differ in covariate composition even when trajectory shapes are similar, because covariates directly enter the allocation probabilities [2203.12280]. Single-atom spatio-temporal DDPs retain a positive lower bound on dependence unless the atoms vary as well [2303.17177]. Logistic stick-breaking for multinomial probabilities is asymmetric, so category ordering matters [1506.05843]. Balanced finite trees mitigate several lopsided-stick pathologies, but infinite balanced-tree constructions are delicate and can degenerate without depth-dependent regularization [2208.02806].

Taken together, these results establish location-dependent stick-breaking priors as a broad methodological class rather than a single model. Their unifying principle is simple: retain the stick-breaking normalization, but let the break probabilities vary with an external index. The resulting family encompasses predictor-dependent density regression, covariate-dependent longitudinal clustering, spatial and spatio-temporal DDPs, tree-structured gating models, and finite multinomial regressions. The primary modeling choice is therefore not whether to use stick-breaking, but where to place the dependence—in the weights, the atoms, or both—and how strongly to couple nearby locations through that dependence [2203.12280; 2303.17177; 2208.02806].

Source: https://www.emergentmind.com/topics/location-dependent-stick-breaking-prior