---
title: Markov Switching Dynamic Shrinkage Process
url: https://www.emergentmind.com/topics/markov-switching-dynamic-shrinkage-process-msdsp
type: topic
---

# Markov Switching Dynamic Shrinkage Process

The Markov Switching Dynamic Shrinkage Process (MSDSP) is a Bayesian prior and state-process construction for time-varying parameter models in which each coefficient is governed by two latent mechanisms: a dynamically shrunk “shadow” coefficient and a two-state Markov indicator that determines whether the realized coefficient is active or exactly zero. In the formulation introduced for exchange-rate prediction, MSDSP nests the Dynamic Shrinkage Process (DSP) of Kowal et al. and is intended to represent three behaviors within one hierarchical model: exact sparsity, local constancy, and both sudden and gradual parameter movement [2507.14408].

## 1. Concept and intellectual lineage

MSDSP arises in the standard time-varying parameter setting
\[
y_t = \bm{x}'_t\bm{\beta}_t +\epsilon_t,\qquad \epsilon_t \sim N\left(0,\sigma_t^2 \right),
\]
with coefficient evolution
\[
\bm{\beta}_t = \bm{\beta}_{t-1}+\bm{\omega}_t,\qquad \bm{\omega}_t \sim N\left(0, \Omega_t\right),
\]
and, in the macro-finance applications motivating the model, stochastic volatility in the observation equation,
\[
g_t = \mu_g+\phi_g (g_{t-1}-\mu_g)+e_t,\qquad e_t \sim N\left(0,\sigma_g^2\right), \quad \sigma_t^2 = \exp(g_t).
\]

The immediate antecedent is the DSP of Kowal et al., which models the log-variance of state innovations through a latent AR(1) process with \(Z\)-distributed shocks. In its scalar form,
\[
\tilde{\beta}_{it} = \tilde{\beta}_{i,t-1} + \omega_{it},\qquad \omega_{it}\sim N(0,\exp(h_{it})),
\]
\[
h_{i,t+1}= \mu_i + \phi_i (h_{it} -\mu_i) + \eta_{i,t+1},\qquad \eta_{i,t+1}\sim Z(\alpha_h, \beta_h, 0, 1).
\]
The critical DSP mechanism is that very negative shocks to \(h_{it}\) push \(\exp(h_{it})\) toward zero, so the innovation variance collapses and the coefficient becomes locally constant. DSP therefore shrinks coefficient *changes* toward zero, not the coefficient itself toward zero [1707.00763].

MSDSP adds a discrete inclusion layer to that continuous shrinkage mechanism. This addition addresses two limitations stated explicitly for DSP: it does not set coefficients exactly to zero, and it does not explicitly separate active from inactive periods of a regressor. The Markov-switching layer provides that separation by letting each coefficient switch between a DSP-governed active state and an exactly zero state [2507.14408].

## 2. Hierarchical formulation

The core masking equation is
\[
\bm{\beta}_{t} = \bm{s}_t\circ \tilde{\bm{\beta}}_t + (\bm{\ell}-\bm{s}_t)\circ \bm{0},
\]
where \(\bm{s}_t=(s_{1t},\dots,s_{pt})'\in\{0,1\}^p\), \(\tilde{\bm{\beta}}_t\) is the vector of shadow coefficients, \(\bm{\ell}\) is a vector of ones, and \(\circ\) denotes the Hadamard product. Componentwise,
\[
\beta_{it} =
\begin{cases}
\tilde{\beta}_{it}, & s_{it}=1,\\
0, & s_{it}=0.
\end{cases}
\]
This is the exact source of sparsity: \(s_{it}=0\) implies \(\beta_{it}=0\).

The shadow coefficients evolve under DSP dynamics,
\[
\tilde{\beta}_{it} =\tilde{\beta}_{i,t-1} + \omega_{it},\qquad \omega_{it} \sim N\left(0, \exp(h_{it})\right),
\]
\[
h_{i,t+1} = \mu_i + \phi_i (h_{it} -\mu_i) + \eta_{i,t+1},\qquad \eta_{i,t+1}\sim Z(\alpha_h, \beta_h, 0, 1).
\]
Here \(\mu_i\) and \(\phi_i\) determine the level and persistence of the innovation log-variance, while the \(Z\)-distribution is introduced as a flexible normal mixture obtained by logistic transformation of a Beta\((\alpha_h,\beta_h)\) variable and allows heavy tails and skewness.

The on/off mechanism is a collection of independent two-state Markov chains,
\[
P(s_{it}=k\mid s_{i,t-1}=j) = P^i_{jk},\qquad j,k\in\{0,1\},
\]
with one transition matrix \(P^i\in\mathbb{R}^{2\times2}\) per regressor. State \(0\) is the zero-state and state \(1\) is the DSP-state. Persistence in \(P^i\) yields persistence in coefficient inclusion, so coefficients may switch on, stay active, and later switch off.

In the observation layer, MSDSP is embedded in a standard univariate DLM with stochastic volatility,
\[
y_t = \bm{x}'_t\bm{\beta}_t +\epsilon_t,\qquad \epsilon_t \sim N\left(0,\exp(g_t)\right),
\]
\[
g_t = \mu_g+\phi_g (g_{t-1}-\mu_g)+e_t,\qquad e_t \sim N\left(0,\sigma_g^2\right).
\]
The construction is not restricted to regression coefficients. The same masking-plus-DSP architecture is also placed on dynamic combination weights in Bayesian predictive synthesis, where the latent weights may be shut off exactly or evolve smoothly when active [2507.14408].

## 3. Shrinkage behavior, interpretation, and nested models

MSDSP was introduced to combine three behaviors in one prior specification. First, exact sparsity is achieved through the zero-state \(s_{it}=0\). Second, approximate constancy is achieved in the active state when the DSP drives \(\exp(h_{it})\) near zero, suppressing innovations in \(\tilde{\beta}_{it}\). Third, both abrupt and gradual changes are represented: abrupt changes through regime switches \(0\leftrightarrow 1\), and gradual changes through the random-walk evolution of active shadow coefficients [2507.14408].

A common misunderstanding is to treat DSP itself as a variable-selection prior. It is not. DSP suppresses *changes* in coefficients by shrinking state innovation variances; it does not create exact zeros in coefficients. MSDSP adds exact sparsity by introducing the Markov mask \(s_{it}\). Conversely, MSDSP is not merely a switching regression with piecewise-constant coefficients: when \(s_{it}=1\), the coefficient remains fully dynamic because the shadow state continues to evolve under DSP.

The model nests several standard cases. If the Markov chain degenerates so that \(s_{it}=1\) for all \(t\), then \(\bm{\beta}_t=\tilde{\bm{\beta}}_t\) and the model reduces to the DSP-based TVP specification of Kowal et al. If \(\exp(h_{it})\) is constant over time and \(s_{it}=1\) always, the model becomes a standard random-walk TVP model. If \(\tilde{\beta}_{it}\) is fixed but \(s_{it}\) remains Markov, the model becomes a dynamic SSVS-like sparsity specification. If \(\omega_{it}\equiv 0\) and \(s_{it}\equiv 1\), the model collapses to static constant-coefficient regression [2507.14408].

Within its DSP core, MSDSP inherits the logic of dynamic global-local shrinkage. In Kowal et al., local innovation scales are driven by an AR(1) log-variance process,
\[
h_{t+1} = \mu + \phi(h_t-\mu)+\eta_t,\qquad \eta_t\sim Z(\alpha,\beta,0,1),
\]
which induces dependence in the shrinkage parameter
\[
\kappa_t = \frac{1}{1+\tau^2\lambda_t^2}.
\]
For the dynamic horseshoe case, the theory shows persistence of low-shrinkage and high-shrinkage episodes across adjacent times, yielding clustered runs of minimal or aggressive shrinkage [1707.00763]. Because MSDSP embeds DSP in its active state, a plausible implication is that active periods inherit this temporal clustering of local flexibility, while the Markov mask adds a second, discrete layer of persistence in coefficient relevance.

## 4. Posterior inference and computation

Posterior inference targets the joint distribution of shadow coefficients, inclusion states, observation volatilities, DSP log-volatilities, Markov transition matrices, and auxiliary variables for the Polya-Gamma and mixture-of-normals representations. The unknown set \(\Psi\) contains, in the notation of the model, \(\tilde{\bm{\beta}}\), \(\bm{s}\), \(g_{1:T}\), \((\mu_g,\phi_g,\sigma_g^2)\), \(H=(\bm{h}_1,\dots,\bm{h}_T)\), \((\bm{\mu},\bm{\phi})\), \(P^i\), and the corresponding augmentation variables [2507.14408].

The computational strategy combines three standard ingredients. First, DSP sampling follows Kowal et al.’s Polya-Gamma augmentation for the \(Z\)-distributed innovation terms, which yields conditional Gaussian updates for the latent log-variances [1707.00763]. Second, observation stochastic volatility is handled through the Omori et al. finite mixture-of-normals approximation. Third, both the Gaussian state blocks and the discrete Markov chains are sampled by forward-filtering backward-sampling.

The Gibbs-type algorithm described for MSDSP proceeds as follows. Shadow coefficient paths \(\tilde{\bm{\beta}}_{1:T}\) are sampled conditional on \(\bm{y},\bm{X},\bm{s},\bm{g}\), and \(H\) using FFBS. Inclusion states \(\bm{s}\) are sampled coefficientwise as 2-state hidden Markov models, again with FFBS. The realized coefficients \(\bm{\beta}_t = \bm{s}_t\circ \tilde{\bm{\beta}}_t\) are then updated deterministically. Observation volatilities \(g_{1:T}\) are sampled through the stochastic-volatility mixture representation. Transition matrices \(P^i\) are updated by Dirichlet/posterior-Dirichlet steps using counts of \(0\to 0\), \(0\to 1\), \(1\to 0\), and \(1\to 1\) transitions. The DSP volatility block \(H\) and its parameters are updated using Kowal et al.’s sampling strategy, including auxiliary \(\bm{\xi}\) and \(\bm{\xi}_\mu\) variables from Polya-Gamma-like augmentation [2507.14408].

The implementation remarks are equally important. The algorithm decomposes across coefficients for the DSP and Markov-chain blocks, so it is embarrassingly parallel. The computational bottleneck is sampling \(p\) shadow coefficient trajectories and \(p\) Markov chains, but the paper notes that in typical macro applications \(p\) is small. In the reported implementation, the sampler uses 30,000 burn-in iterations and retains 4,000 posterior draws with thinning by 5 [2507.14408].

## 5. Applications in exchange-rate prediction and Bayesian predictive synthesis

The empirical motivation for MSDSP is the Meese–Rogoff puzzle. The response variable is the exchange-rate change
\[
y_t = e_t - e_{t-1},
\]
with \(e_t\) the log spot exchange rate. The benchmark is a random walk with stochastic volatility and, alternatively, a random walk with stochastic volatility and no drift. Economic models are embedded in the MSDSP TVP regression,
\[
y_{t+h} = \bm{x}'_t \bm{\beta}_t + \epsilon_t,\qquad \epsilon_t\sim N(0,\exp(g_t)),
\]
using direct \(h\)-step-ahead regressions. The competing economic specifications include Interest Rate Parity, Taylor Rule, Purchasing Power Parity, Monetary Model, and commodity-price models based on oil, gold, and copper [2507.14408].

Predictive evaluation is based on LPDR relative to the RW-SV benchmark, CRPS, RMSFE, tail coverage rates at lower and upper predictive quantiles, and Model Confidence Set analysis. The reported findings are sharply differentiated across model classes. Constant-parameter linear economic models generally perform worse than RW-SV. DSP models without Markov switching improve upon linear models but still often underperform RW-SV. MSDSP versions of the same economic models systematically outperform RW-SV: they have positive LPDR across horizons, often greater than 3, lower CRPS, smaller RMSFE, and tail coverage closer to nominal coverage, especially for upper tails. PPP-MSDSP is highlighted as particularly strong across horizons, and MCS rankings place MSDSP economic models at the top while RW-SV remains in the confidence set with lower ranks [2507.14408].

The interpretation advanced in that application is methodological rather than purely empirical. The inferior performance of constant-parameter economic models is not taken as evidence that the underlying economic relations are absent; rather, the claim is that rigid econometric specification obscures relations that are intermittent, sparse, and regime-dependent. MSDSP is designed precisely for such intermittent relevance because a coefficient can be zero, locally constant, gradually drifting, or abruptly reactivated.

The same logic is transferred to Bayesian predictive synthesis. In the original DBPS framework,
\[
y_t = \omega_{t, 0}+\sum_{j=1}^L \omega_{t, j} \hat{y}_{t, j}+e_t,\qquad e_t \sim N\left(0, v_t\right),
\]
with dynamic synthesis weights \(\omega_{t,j}\). The MSDSP variant keeps the dynamic combination architecture but places an MSDSP prior on each weight \(\omega_{t,j}\), including the intercept, while the synthesis-error variance follows stochastic volatility:
\[
y_t = \omega_{t, 0}+\sum_{j=1}^L \omega_{t, j} \hat{y}_{t, j}+e_t,\qquad e_t \sim N\left(0, e^{v_t}\right).
\]
The resulting assembly model allows some forecast sources to be shut off exactly, others to remain active with nearly constant weights, and others to adapt over time. Empirically, MSDSP assembly outperforms DSP assembly and original DBPS in all reported metrics and horizons; relative to RW-SV, it is slightly worse at very short horizons but becomes better at longer horizons, with RMSFE better than RW-SV for \(h>9\) [2507.14408].

## 6. Relation to adjacent Markov-switching shrinkage frameworks

MSDSP belongs to a broader family of Bayesian methods that combine dynamic shrinkage with latent regime structure, but neighboring models differ materially in what is switched and what is shrunk. A useful comparison point is the dynamic regression model with Markov switching priors on coefficient variances proposed in “Dynamic sparsity on dynamic regression models,” where each variance takes spike or slab values \(K_{jt}\tau_j^2\) under a two-state Markov chain. That construction induces time-varying sparsity and can be interpreted as a dynamic spike-and-slab prior, but it does not use the DSP shadow-state mechanism that characterizes MSDSP [2009.14131].

Another nearby line is the “Bayesian Dynamic Fused LASSO,” which defines a stationary Markov process with transition density
\[
p(x_t \mid x_{t-1}) \propto \exp\left\{-\alpha |x_t| - \beta |x_t - x_{t-1}|\right\}.
\]
This process combines shrinkage to zero and shrinkage to the previous state within a conditionally Gaussian DLM. Its latent counts \(n_t\) create implicit switching between stronger and weaker shrinkage to zero, but the switching is not an explicit finite-state Markov inclusion layer of the type used in MSDSP [1905.12275].

The dynamic triple gamma prior provides a different form of dynamic shrinkage: a continuous Markov process on innovation variances \(\psi_{jt}\) with triple-gamma marginals and horseshoe as a special case when \(a_j=c_j=\tfrac12\). That framework already exhibits persistence in high- and low-variance phases through a Markov dependence parameter \(\rho_j\), and the paper explicitly states that an MSDSP-like extension would arise by adding a discrete regime indicator \(S_t\) that changes \((a_j,c_j,\rho_j,\omega_j)\) across time [2312.10487].

A further related contribution studies shrinkage estimation under hidden Markov dependence by combining Tweedie-based shrinkage with posterior state probabilities from an HMM. There, the effective shrinkage rule is state-specific and varies along the sequence according to inferred latent states. This is a Markov-modulated shrinkage mechanism, but it is an empirical-Bayes estimator for dependent normal means rather than a DSP-based state evolution prior for TVP coefficients [2003.01873].

These comparisons clarify the distinctive position of MSDSP. Its defining feature is not merely Markovian dependence, nor merely dynamic shrinkage, nor merely exact zeros. It is the combination of a DSP-governed active state and a discrete Markov zero-state, yielding a model in which inclusion, local constancy, gradual drift, and abrupt structural change are represented simultaneously within a single hierarchical specification [2507.14408].

Source: https://www.emergentmind.com/topics/markov-switching-dynamic-shrinkage-process-msdsp