Papers
Topics
Authors
Recent
Search
2000 character limit reached

Markov Switching Dynamic Shrinkage Process

Updated 6 July 2026
  • MSDSP is a Bayesian prior that uses a dynamic shrinkage process and a two-state Markov indicator to model time-varying sparsity and parameter shifts.
  • It nests the dynamic shrinkage process from Kowal et al. to resolve limitations by enabling coefficients to switch between exact zero and active states.
  • The hierarchical model leverages efficient FFBS sampling for posterior inference, making it applicable to exchange-rate prediction and Bayesian predictive synthesis.

The Markov Switching Dynamic Shrinkage Process (MSDSP) is a Bayesian prior and state-process construction for time-varying parameter models in which each coefficient is governed by two latent mechanisms: a dynamically shrunk “shadow” coefficient and a two-state Markov indicator that determines whether the realized coefficient is active or exactly zero. In the formulation introduced for exchange-rate prediction, MSDSP nests the Dynamic Shrinkage Process (DSP) of Kowal et al. and is intended to represent three behaviors within one hierarchical model: exact sparsity, local constancy, and both sudden and gradual parameter movement (Fan et al., 18 Jul 2025).

1. Concept and intellectual lineage

MSDSP arises in the standard time-varying parameter setting

yt=xtβt+ϵt,ϵtN(0,σt2),y_t = \bm{x}'_t\bm{\beta}_t +\epsilon_t,\qquad \epsilon_t \sim N\left(0,\sigma_t^2 \right),

with coefficient evolution

βt=βt1+ωt,ωtN(0,Ωt),\bm{\beta}_t = \bm{\beta}_{t-1}+\bm{\omega}_t,\qquad \bm{\omega}_t \sim N\left(0, \Omega_t\right),

and, in the macro-finance applications motivating the model, stochastic volatility in the observation equation,

gt=μg+ϕg(gt1μg)+et,etN(0,σg2),σt2=exp(gt).g_t = \mu_g+\phi_g (g_{t-1}-\mu_g)+e_t,\qquad e_t \sim N\left(0,\sigma_g^2\right), \quad \sigma_t^2 = \exp(g_t).

The immediate antecedent is the DSP of Kowal et al., which models the log-variance of state innovations through a latent AR(1) process with ZZ-distributed shocks. In its scalar form,

β~it=β~i,t1+ωit,ωitN(0,exp(hit)),\tilde{\beta}_{it} = \tilde{\beta}_{i,t-1} + \omega_{it},\qquad \omega_{it}\sim N(0,\exp(h_{it})),

hi,t+1=μi+ϕi(hitμi)+ηi,t+1,ηi,t+1Z(αh,βh,0,1).h_{i,t+1}= \mu_i + \phi_i (h_{it} -\mu_i) + \eta_{i,t+1},\qquad \eta_{i,t+1}\sim Z(\alpha_h, \beta_h, 0, 1).

The critical DSP mechanism is that very negative shocks to hith_{it} push exp(hit)\exp(h_{it}) toward zero, so the innovation variance collapses and the coefficient becomes locally constant. DSP therefore shrinks coefficient changes toward zero, not the coefficient itself toward zero (Kowal et al., 2017).

MSDSP adds a discrete inclusion layer to that continuous shrinkage mechanism. This addition addresses two limitations stated explicitly for DSP: it does not set coefficients exactly to zero, and it does not explicitly separate active from inactive periods of a regressor. The Markov-switching layer provides that separation by letting each coefficient switch between a DSP-governed active state and an exactly zero state (Fan et al., 18 Jul 2025).

2. Hierarchical formulation

The core masking equation is

βt=stβ~t+(st)0,\bm{\beta}_{t} = \bm{s}_t\circ \tilde{\bm{\beta}}_t + (\bm{\ell}-\bm{s}_t)\circ \bm{0},

where st=(s1t,,spt){0,1}p\bm{s}_t=(s_{1t},\dots,s_{pt})'\in\{0,1\}^p, βt=βt1+ωt,ωtN(0,Ωt),\bm{\beta}_t = \bm{\beta}_{t-1}+\bm{\omega}_t,\qquad \bm{\omega}_t \sim N\left(0, \Omega_t\right),0 is the vector of shadow coefficients, βt=βt1+ωt,ωtN(0,Ωt),\bm{\beta}_t = \bm{\beta}_{t-1}+\bm{\omega}_t,\qquad \bm{\omega}_t \sim N\left(0, \Omega_t\right),1 is a vector of ones, and βt=βt1+ωt,ωtN(0,Ωt),\bm{\beta}_t = \bm{\beta}_{t-1}+\bm{\omega}_t,\qquad \bm{\omega}_t \sim N\left(0, \Omega_t\right),2 denotes the Hadamard product. Componentwise,

βt=βt1+ωt,ωtN(0,Ωt),\bm{\beta}_t = \bm{\beta}_{t-1}+\bm{\omega}_t,\qquad \bm{\omega}_t \sim N\left(0, \Omega_t\right),3

This is the exact source of sparsity: βt=βt1+ωt,ωtN(0,Ωt),\bm{\beta}_t = \bm{\beta}_{t-1}+\bm{\omega}_t,\qquad \bm{\omega}_t \sim N\left(0, \Omega_t\right),4 implies βt=βt1+ωt,ωtN(0,Ωt),\bm{\beta}_t = \bm{\beta}_{t-1}+\bm{\omega}_t,\qquad \bm{\omega}_t \sim N\left(0, \Omega_t\right),5.

The shadow coefficients evolve under DSP dynamics,

βt=βt1+ωt,ωtN(0,Ωt),\bm{\beta}_t = \bm{\beta}_{t-1}+\bm{\omega}_t,\qquad \bm{\omega}_t \sim N\left(0, \Omega_t\right),6

βt=βt1+ωt,ωtN(0,Ωt),\bm{\beta}_t = \bm{\beta}_{t-1}+\bm{\omega}_t,\qquad \bm{\omega}_t \sim N\left(0, \Omega_t\right),7

Here βt=βt1+ωt,ωtN(0,Ωt),\bm{\beta}_t = \bm{\beta}_{t-1}+\bm{\omega}_t,\qquad \bm{\omega}_t \sim N\left(0, \Omega_t\right),8 and βt=βt1+ωt,ωtN(0,Ωt),\bm{\beta}_t = \bm{\beta}_{t-1}+\bm{\omega}_t,\qquad \bm{\omega}_t \sim N\left(0, \Omega_t\right),9 determine the level and persistence of the innovation log-variance, while the gt=μg+ϕg(gt1μg)+et,etN(0,σg2),σt2=exp(gt).g_t = \mu_g+\phi_g (g_{t-1}-\mu_g)+e_t,\qquad e_t \sim N\left(0,\sigma_g^2\right), \quad \sigma_t^2 = \exp(g_t).0-distribution is introduced as a flexible normal mixture obtained by logistic transformation of a Betagt=μg+ϕg(gt1μg)+et,etN(0,σg2),σt2=exp(gt).g_t = \mu_g+\phi_g (g_{t-1}-\mu_g)+e_t,\qquad e_t \sim N\left(0,\sigma_g^2\right), \quad \sigma_t^2 = \exp(g_t).1 variable and allows heavy tails and skewness.

The on/off mechanism is a collection of independent two-state Markov chains,

gt=μg+ϕg(gt1μg)+et,etN(0,σg2),σt2=exp(gt).g_t = \mu_g+\phi_g (g_{t-1}-\mu_g)+e_t,\qquad e_t \sim N\left(0,\sigma_g^2\right), \quad \sigma_t^2 = \exp(g_t).2

with one transition matrix gt=μg+ϕg(gt1μg)+et,etN(0,σg2),σt2=exp(gt).g_t = \mu_g+\phi_g (g_{t-1}-\mu_g)+e_t,\qquad e_t \sim N\left(0,\sigma_g^2\right), \quad \sigma_t^2 = \exp(g_t).3 per regressor. State gt=μg+ϕg(gt1μg)+et,etN(0,σg2),σt2=exp(gt).g_t = \mu_g+\phi_g (g_{t-1}-\mu_g)+e_t,\qquad e_t \sim N\left(0,\sigma_g^2\right), \quad \sigma_t^2 = \exp(g_t).4 is the zero-state and state gt=μg+ϕg(gt1μg)+et,etN(0,σg2),σt2=exp(gt).g_t = \mu_g+\phi_g (g_{t-1}-\mu_g)+e_t,\qquad e_t \sim N\left(0,\sigma_g^2\right), \quad \sigma_t^2 = \exp(g_t).5 is the DSP-state. Persistence in gt=μg+ϕg(gt1μg)+et,etN(0,σg2),σt2=exp(gt).g_t = \mu_g+\phi_g (g_{t-1}-\mu_g)+e_t,\qquad e_t \sim N\left(0,\sigma_g^2\right), \quad \sigma_t^2 = \exp(g_t).6 yields persistence in coefficient inclusion, so coefficients may switch on, stay active, and later switch off.

In the observation layer, MSDSP is embedded in a standard univariate DLM with stochastic volatility,

gt=μg+ϕg(gt1μg)+et,etN(0,σg2),σt2=exp(gt).g_t = \mu_g+\phi_g (g_{t-1}-\mu_g)+e_t,\qquad e_t \sim N\left(0,\sigma_g^2\right), \quad \sigma_t^2 = \exp(g_t).7

gt=μg+ϕg(gt1μg)+et,etN(0,σg2),σt2=exp(gt).g_t = \mu_g+\phi_g (g_{t-1}-\mu_g)+e_t,\qquad e_t \sim N\left(0,\sigma_g^2\right), \quad \sigma_t^2 = \exp(g_t).8

The construction is not restricted to regression coefficients. The same masking-plus-DSP architecture is also placed on dynamic combination weights in Bayesian predictive synthesis, where the latent weights may be shut off exactly or evolve smoothly when active (Fan et al., 18 Jul 2025).

3. Shrinkage behavior, interpretation, and nested models

MSDSP was introduced to combine three behaviors in one prior specification. First, exact sparsity is achieved through the zero-state gt=μg+ϕg(gt1μg)+et,etN(0,σg2),σt2=exp(gt).g_t = \mu_g+\phi_g (g_{t-1}-\mu_g)+e_t,\qquad e_t \sim N\left(0,\sigma_g^2\right), \quad \sigma_t^2 = \exp(g_t).9. Second, approximate constancy is achieved in the active state when the DSP drives ZZ0 near zero, suppressing innovations in ZZ1. Third, both abrupt and gradual changes are represented: abrupt changes through regime switches ZZ2, and gradual changes through the random-walk evolution of active shadow coefficients (Fan et al., 18 Jul 2025).

A common misunderstanding is to treat DSP itself as a variable-selection prior. It is not. DSP suppresses changes in coefficients by shrinking state innovation variances; it does not create exact zeros in coefficients. MSDSP adds exact sparsity by introducing the Markov mask ZZ3. Conversely, MSDSP is not merely a switching regression with piecewise-constant coefficients: when ZZ4, the coefficient remains fully dynamic because the shadow state continues to evolve under DSP.

The model nests several standard cases. If the Markov chain degenerates so that ZZ5 for all ZZ6, then ZZ7 and the model reduces to the DSP-based TVP specification of Kowal et al. If ZZ8 is constant over time and ZZ9 always, the model becomes a standard random-walk TVP model. If β~it=β~i,t1+ωit,ωitN(0,exp(hit)),\tilde{\beta}_{it} = \tilde{\beta}_{i,t-1} + \omega_{it},\qquad \omega_{it}\sim N(0,\exp(h_{it})),0 is fixed but β~it=β~i,t1+ωit,ωitN(0,exp(hit)),\tilde{\beta}_{it} = \tilde{\beta}_{i,t-1} + \omega_{it},\qquad \omega_{it}\sim N(0,\exp(h_{it})),1 remains Markov, the model becomes a dynamic SSVS-like sparsity specification. If β~it=β~i,t1+ωit,ωitN(0,exp(hit)),\tilde{\beta}_{it} = \tilde{\beta}_{i,t-1} + \omega_{it},\qquad \omega_{it}\sim N(0,\exp(h_{it})),2 and β~it=β~i,t1+ωit,ωitN(0,exp(hit)),\tilde{\beta}_{it} = \tilde{\beta}_{i,t-1} + \omega_{it},\qquad \omega_{it}\sim N(0,\exp(h_{it})),3, the model collapses to static constant-coefficient regression (Fan et al., 18 Jul 2025).

Within its DSP core, MSDSP inherits the logic of dynamic global-local shrinkage. In Kowal et al., local innovation scales are driven by an AR(1) log-variance process,

β~it=β~i,t1+ωit,ωitN(0,exp(hit)),\tilde{\beta}_{it} = \tilde{\beta}_{i,t-1} + \omega_{it},\qquad \omega_{it}\sim N(0,\exp(h_{it})),4

which induces dependence in the shrinkage parameter

β~it=β~i,t1+ωit,ωitN(0,exp(hit)),\tilde{\beta}_{it} = \tilde{\beta}_{i,t-1} + \omega_{it},\qquad \omega_{it}\sim N(0,\exp(h_{it})),5

For the dynamic horseshoe case, the theory shows persistence of low-shrinkage and high-shrinkage episodes across adjacent times, yielding clustered runs of minimal or aggressive shrinkage (Kowal et al., 2017). Because MSDSP embeds DSP in its active state, a plausible implication is that active periods inherit this temporal clustering of local flexibility, while the Markov mask adds a second, discrete layer of persistence in coefficient relevance.

4. Posterior inference and computation

Posterior inference targets the joint distribution of shadow coefficients, inclusion states, observation volatilities, DSP log-volatilities, Markov transition matrices, and auxiliary variables for the Polya-Gamma and mixture-of-normals representations. The unknown set β~it=β~i,t1+ωit,ωitN(0,exp(hit)),\tilde{\beta}_{it} = \tilde{\beta}_{i,t-1} + \omega_{it},\qquad \omega_{it}\sim N(0,\exp(h_{it})),6 contains, in the notation of the model, β~it=β~i,t1+ωit,ωitN(0,exp(hit)),\tilde{\beta}_{it} = \tilde{\beta}_{i,t-1} + \omega_{it},\qquad \omega_{it}\sim N(0,\exp(h_{it})),7, β~it=β~i,t1+ωit,ωitN(0,exp(hit)),\tilde{\beta}_{it} = \tilde{\beta}_{i,t-1} + \omega_{it},\qquad \omega_{it}\sim N(0,\exp(h_{it})),8, β~it=β~i,t1+ωit,ωitN(0,exp(hit)),\tilde{\beta}_{it} = \tilde{\beta}_{i,t-1} + \omega_{it},\qquad \omega_{it}\sim N(0,\exp(h_{it})),9, hi,t+1=μi+ϕi(hitμi)+ηi,t+1,ηi,t+1Z(αh,βh,0,1).h_{i,t+1}= \mu_i + \phi_i (h_{it} -\mu_i) + \eta_{i,t+1},\qquad \eta_{i,t+1}\sim Z(\alpha_h, \beta_h, 0, 1).0, hi,t+1=μi+ϕi(hitμi)+ηi,t+1,ηi,t+1Z(αh,βh,0,1).h_{i,t+1}= \mu_i + \phi_i (h_{it} -\mu_i) + \eta_{i,t+1},\qquad \eta_{i,t+1}\sim Z(\alpha_h, \beta_h, 0, 1).1, hi,t+1=μi+ϕi(hitμi)+ηi,t+1,ηi,t+1Z(αh,βh,0,1).h_{i,t+1}= \mu_i + \phi_i (h_{it} -\mu_i) + \eta_{i,t+1},\qquad \eta_{i,t+1}\sim Z(\alpha_h, \beta_h, 0, 1).2, hi,t+1=μi+ϕi(hitμi)+ηi,t+1,ηi,t+1Z(αh,βh,0,1).h_{i,t+1}= \mu_i + \phi_i (h_{it} -\mu_i) + \eta_{i,t+1},\qquad \eta_{i,t+1}\sim Z(\alpha_h, \beta_h, 0, 1).3, and the corresponding augmentation variables (Fan et al., 18 Jul 2025).

The computational strategy combines three standard ingredients. First, DSP sampling follows Kowal et al.’s Polya-Gamma augmentation for the hi,t+1=μi+ϕi(hitμi)+ηi,t+1,ηi,t+1Z(αh,βh,0,1).h_{i,t+1}= \mu_i + \phi_i (h_{it} -\mu_i) + \eta_{i,t+1},\qquad \eta_{i,t+1}\sim Z(\alpha_h, \beta_h, 0, 1).4-distributed innovation terms, which yields conditional Gaussian updates for the latent log-variances (Kowal et al., 2017). Second, observation stochastic volatility is handled through the Omori et al. finite mixture-of-normals approximation. Third, both the Gaussian state blocks and the discrete Markov chains are sampled by forward-filtering backward-sampling.

The Gibbs-type algorithm described for MSDSP proceeds as follows. Shadow coefficient paths hi,t+1=μi+ϕi(hitμi)+ηi,t+1,ηi,t+1Z(αh,βh,0,1).h_{i,t+1}= \mu_i + \phi_i (h_{it} -\mu_i) + \eta_{i,t+1},\qquad \eta_{i,t+1}\sim Z(\alpha_h, \beta_h, 0, 1).5 are sampled conditional on hi,t+1=μi+ϕi(hitμi)+ηi,t+1,ηi,t+1Z(αh,βh,0,1).h_{i,t+1}= \mu_i + \phi_i (h_{it} -\mu_i) + \eta_{i,t+1},\qquad \eta_{i,t+1}\sim Z(\alpha_h, \beta_h, 0, 1).6, and hi,t+1=μi+ϕi(hitμi)+ηi,t+1,ηi,t+1Z(αh,βh,0,1).h_{i,t+1}= \mu_i + \phi_i (h_{it} -\mu_i) + \eta_{i,t+1},\qquad \eta_{i,t+1}\sim Z(\alpha_h, \beta_h, 0, 1).7 using FFBS. Inclusion states hi,t+1=μi+ϕi(hitμi)+ηi,t+1,ηi,t+1Z(αh,βh,0,1).h_{i,t+1}= \mu_i + \phi_i (h_{it} -\mu_i) + \eta_{i,t+1},\qquad \eta_{i,t+1}\sim Z(\alpha_h, \beta_h, 0, 1).8 are sampled coefficientwise as 2-state hidden Markov models, again with FFBS. The realized coefficients hi,t+1=μi+ϕi(hitμi)+ηi,t+1,ηi,t+1Z(αh,βh,0,1).h_{i,t+1}= \mu_i + \phi_i (h_{it} -\mu_i) + \eta_{i,t+1},\qquad \eta_{i,t+1}\sim Z(\alpha_h, \beta_h, 0, 1).9 are then updated deterministically. Observation volatilities hith_{it}0 are sampled through the stochastic-volatility mixture representation. Transition matrices hith_{it}1 are updated by Dirichlet/posterior-Dirichlet steps using counts of hith_{it}2, hith_{it}3, hith_{it}4, and hith_{it}5 transitions. The DSP volatility block hith_{it}6 and its parameters are updated using Kowal et al.’s sampling strategy, including auxiliary hith_{it}7 and hith_{it}8 variables from Polya-Gamma-like augmentation (Fan et al., 18 Jul 2025).

The implementation remarks are equally important. The algorithm decomposes across coefficients for the DSP and Markov-chain blocks, so it is embarrassingly parallel. The computational bottleneck is sampling hith_{it}9 shadow coefficient trajectories and exp(hit)\exp(h_{it})0 Markov chains, but the paper notes that in typical macro applications exp(hit)\exp(h_{it})1 is small. In the reported implementation, the sampler uses 30,000 burn-in iterations and retains 4,000 posterior draws with thinning by 5 (Fan et al., 18 Jul 2025).

5. Applications in exchange-rate prediction and Bayesian predictive synthesis

The empirical motivation for MSDSP is the Meese–Rogoff puzzle. The response variable is the exchange-rate change

exp(hit)\exp(h_{it})2

with exp(hit)\exp(h_{it})3 the log spot exchange rate. The benchmark is a random walk with stochastic volatility and, alternatively, a random walk with stochastic volatility and no drift. Economic models are embedded in the MSDSP TVP regression,

exp(hit)\exp(h_{it})4

using direct exp(hit)\exp(h_{it})5-step-ahead regressions. The competing economic specifications include Interest Rate Parity, Taylor Rule, Purchasing Power Parity, Monetary Model, and commodity-price models based on oil, gold, and copper (Fan et al., 18 Jul 2025).

Predictive evaluation is based on LPDR relative to the RW-SV benchmark, CRPS, RMSFE, tail coverage rates at lower and upper predictive quantiles, and Model Confidence Set analysis. The reported findings are sharply differentiated across model classes. Constant-parameter linear economic models generally perform worse than RW-SV. DSP models without Markov switching improve upon linear models but still often underperform RW-SV. MSDSP versions of the same economic models systematically outperform RW-SV: they have positive LPDR across horizons, often greater than 3, lower CRPS, smaller RMSFE, and tail coverage closer to nominal coverage, especially for upper tails. PPP-MSDSP is highlighted as particularly strong across horizons, and MCS rankings place MSDSP economic models at the top while RW-SV remains in the confidence set with lower ranks (Fan et al., 18 Jul 2025).

The interpretation advanced in that application is methodological rather than purely empirical. The inferior performance of constant-parameter economic models is not taken as evidence that the underlying economic relations are absent; rather, the claim is that rigid econometric specification obscures relations that are intermittent, sparse, and regime-dependent. MSDSP is designed precisely for such intermittent relevance because a coefficient can be zero, locally constant, gradually drifting, or abruptly reactivated.

The same logic is transferred to Bayesian predictive synthesis. In the original DBPS framework,

exp(hit)\exp(h_{it})6

with dynamic synthesis weights exp(hit)\exp(h_{it})7. The MSDSP variant keeps the dynamic combination architecture but places an MSDSP prior on each weight exp(hit)\exp(h_{it})8, including the intercept, while the synthesis-error variance follows stochastic volatility: exp(hit)\exp(h_{it})9 The resulting assembly model allows some forecast sources to be shut off exactly, others to remain active with nearly constant weights, and others to adapt over time. Empirically, MSDSP assembly outperforms DSP assembly and original DBPS in all reported metrics and horizons; relative to RW-SV, it is slightly worse at very short horizons but becomes better at longer horizons, with RMSFE better than RW-SV for βt=stβ~t+(st)0,\bm{\beta}_{t} = \bm{s}_t\circ \tilde{\bm{\beta}}_t + (\bm{\ell}-\bm{s}_t)\circ \bm{0},0 (Fan et al., 18 Jul 2025).

6. Relation to adjacent Markov-switching shrinkage frameworks

MSDSP belongs to a broader family of Bayesian methods that combine dynamic shrinkage with latent regime structure, but neighboring models differ materially in what is switched and what is shrunk. A useful comparison point is the dynamic regression model with Markov switching priors on coefficient variances proposed in “Dynamic sparsity on dynamic regression models,” where each variance takes spike or slab values βt=stβ~t+(st)0,\bm{\beta}_{t} = \bm{s}_t\circ \tilde{\bm{\beta}}_t + (\bm{\ell}-\bm{s}_t)\circ \bm{0},1 under a two-state Markov chain. That construction induces time-varying sparsity and can be interpreted as a dynamic spike-and-slab prior, but it does not use the DSP shadow-state mechanism that characterizes MSDSP (Uribe et al., 2020).

Another nearby line is the “Bayesian Dynamic Fused LASSO,” which defines a stationary Markov process with transition density

βt=stβ~t+(st)0,\bm{\beta}_{t} = \bm{s}_t\circ \tilde{\bm{\beta}}_t + (\bm{\ell}-\bm{s}_t)\circ \bm{0},2

This process combines shrinkage to zero and shrinkage to the previous state within a conditionally Gaussian DLM. Its latent counts βt=stβ~t+(st)0,\bm{\beta}_{t} = \bm{s}_t\circ \tilde{\bm{\beta}}_t + (\bm{\ell}-\bm{s}_t)\circ \bm{0},3 create implicit switching between stronger and weaker shrinkage to zero, but the switching is not an explicit finite-state Markov inclusion layer of the type used in MSDSP (Irie, 2019).

The dynamic triple gamma prior provides a different form of dynamic shrinkage: a continuous Markov process on innovation variances βt=stβ~t+(st)0,\bm{\beta}_{t} = \bm{s}_t\circ \tilde{\bm{\beta}}_t + (\bm{\ell}-\bm{s}_t)\circ \bm{0},4 with triple-gamma marginals and horseshoe as a special case when βt=stβ~t+(st)0,\bm{\beta}_{t} = \bm{s}_t\circ \tilde{\bm{\beta}}_t + (\bm{\ell}-\bm{s}_t)\circ \bm{0},5. That framework already exhibits persistence in high- and low-variance phases through a Markov dependence parameter βt=stβ~t+(st)0,\bm{\beta}_{t} = \bm{s}_t\circ \tilde{\bm{\beta}}_t + (\bm{\ell}-\bm{s}_t)\circ \bm{0},6, and the paper explicitly states that an MSDSP-like extension would arise by adding a discrete regime indicator βt=stβ~t+(st)0,\bm{\beta}_{t} = \bm{s}_t\circ \tilde{\bm{\beta}}_t + (\bm{\ell}-\bm{s}_t)\circ \bm{0},7 that changes βt=stβ~t+(st)0,\bm{\beta}_{t} = \bm{s}_t\circ \tilde{\bm{\beta}}_t + (\bm{\ell}-\bm{s}_t)\circ \bm{0},8 across time (Knaus et al., 2023).

A further related contribution studies shrinkage estimation under hidden Markov dependence by combining Tweedie-based shrinkage with posterior state probabilities from an HMM. There, the effective shrinkage rule is state-specific and varies along the sequence according to inferred latent states. This is a Markov-modulated shrinkage mechanism, but it is an empirical-Bayes estimator for dependent normal means rather than a DSP-based state evolution prior for TVP coefficients (Gang et al., 2020).

These comparisons clarify the distinctive position of MSDSP. Its defining feature is not merely Markovian dependence, nor merely dynamic shrinkage, nor merely exact zeros. It is the combination of a DSP-governed active state and a discrete Markov zero-state, yielding a model in which inclusion, local constancy, gradual drift, and abrupt structural change are represented simultaneously within a single hierarchical specification (Fan et al., 18 Jul 2025).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Markov Switching Dynamic Shrinkage Process (MSDSP).