---
title: State-Dependent Local Projections
url: https://www.emergentmind.com/topics/state-dependent-local-projections-lps
type: topic
---

# State-Dependent Local Projections

Searching arXiv for the specified papers on state-dependent local projections.
arxiv_search.query({"search_query":"id:2606.13519 OR id:2604.18778 OR id:2601.01622 OR id:2602.14455 OR id:2605.05404","start":0,"max_results":10})
State-dependent local projections (LPs) are local-projection regressions in which the response to a shock is indexed by predetermined observables, typically lagged state variables. They are used to estimate how responses to exogenous aggregate shocks vary as a function of observable state variables, and they can be implemented through linear interactions, regime partitions, nonparametric sieves, or semiparametric moment conditions. Recent work clarifies that the object recovered by a state-dependent LP depends on the specification and assumptions: under minimal exogeneity and smoothness conditions it is a weighted average of causal effects, under conditional linearity in the shock it coincides with a causal state-specific impulse response, and under richer nonlinear designs it can approximate conditional average responses over the joint distribution of shocks and states [2601.01622][2605.05404][2602.14455].

## 1. Formal representations

A general state-dependent LP at horizon \(h\) can be written as
\[
Y_{t+h} = \bigl[f(S_{t-1})\bigr]' X_t \beta^h + \text{error}_{h,t+h},
\]
where \(X_t\) is the shock, \(S_{t-1}\) is the lagged state, and \(f(\cdot)\) is a user-chosen basis for state dependence. In this formulation, the coefficient vector \(\beta^h\) indexes how the horizon-\(h\) response varies with the chosen state basis [2601.01622].

A binary-state structural formulation makes the same object explicit in terms of conditional means. Let \(S_{t-1}\in\{0,1\}\), let
\[
x_t = \phi(z_{t-1}) + \varepsilon_t,\qquad
y_t = \mu(x_t,z_{t-1},\varepsilon_{2t}),
\]
and define
\[
g_{0,h}(x,z,s)\equiv E[y_{t+h}\mid x_t=x,z_{t-1}=z,S_{t-1}=s].
\]
The population impulse response in state \(s\) is
\[
IRF_h(s)
\equiv E[y_{t+h}(\varepsilon_t+\delta)-y_{t+h}(\varepsilon_t)\mid S_{t-1}=s]
= E[g_{0,h}(x_t+\delta,z_{t-1},s)-g_{0,h}(x_t,z_{t-1},s)\mid S_{t-1}=s].
\]
In the linear parametric projection,
\[
Y_{t+h} = \beta_h(S_{t-1})\,\varepsilon_t + \gamma_h(S_{t-1})' z_{t-1} + u_{t+h},
\]
so that \(\beta_h(s)=IRF_h(s)/\delta\) [2606.13519].

In micro-macro panels, the same idea is expressed through a state-dependent slope:
\[
\mathbb{E}[Y_{i,t+h}\mid \mathcal{F}_{t-1},\varepsilon_t]
=
g_h(S_{i,t-1})\,\varepsilon_t + r_{i,h}(\mathcal{F}_{t-1}),
\]
with finite-\(\delta\) impulse response
\[
\mathrm{IRF}_h(\delta\mid s)
=
\mathbb{E}[Y_{i,t+h}(\varepsilon_t+\delta)-Y_{i,t+h}(\varepsilon_t)\mid S_{i,t-1}=s]
=
g_h(s)\,\delta
\]
under linearity in \(\varepsilon_t\) [2605.05404].

A distinct but related formulation is the piecewise-constant or clustered LP:
\[
y_{t+h}
=
\sum_{k=1}^K D_{k,t}\,\beta_k^h\,\varepsilon_t
+
\sum_{k=1}^K D_{k,t}\,\gamma_k^h{}' w_t
+
v_{t+h},
\]
where \(D_{k,t}=\mathbf 1\{s_t\in C_k\}\) indexes a partition of the state space into \(K\) clusters [2604.18778].

These formulations differ in parameterization but share a common objective: estimation of horizon-specific responses conditional on lagged observables. The main differences lie in the estimand, the required assumptions, and the approximation imposed on state heterogeneity.

## 2. Estimands and causal content

Under Assumption sLP, Winkler shows that state-dependent LPs recover weighted averages of causal effects. The conditional regression function
\[
g_h(x,s)=E[Y_{t+h}\mid X_t=x,S_{t-1}=s]
\]
must be locally absolutely continuous in \(x\), and the shock must satisfy
\[
X_t\;\perp\!\!\!\perp\;S_{t-1},
\qquad
X_t\;\perp\!\!\!\perp\;U_{h,t+h}.
\]
Define
\[
\theta_h(s;\omega_X)=\int \omega_X(x)\,\Psi_h'(x,s)\,dx,
\qquad
\omega_X(x)=\frac{\Cov(\mathbf 1\{X_t\ge x\},X_t)}{\Var(X_t)}.
\]
Then the OLS estimand is
\[
\beta^h
=
\bigl(E[f_{t-1}f_{t-1}']\bigr)^{-1}
E\bigl[f_{t-1}\,\theta_h(S_{t-1};\omega_X)\bigr],
\qquad
f_{t-1}:=f(S_{t-1}),
\]
so each component of \(\beta^h\) is a weighted average of state-conditional marginal effects, with \(\omega_X(x)\) depending only on the marginal distribution of \(X_t\) and not on the state [2601.01622].

David et al. obtain a stronger causal interpretation under three high-level conditions: shock exogeneity, predetermined states and controls, and linearity in the shock:
\[
\mathbb{E}[Y_{i,t+h}\mid \mathcal{F}_{t-1},\varepsilon_t]
=
g_h(S_{i,t-1})\,\varepsilon_t+r_{i,h}(\mathcal{F}_{t-1}).
\]
Under these conditions,
\[
g_h(s)
=
\frac{\mathbb{E}[\varepsilon_t Y_{i,t+h}\mid S_{i,t-1}=s]}
{\mathbb{E}[\varepsilon_t^2\mid S_{i,t-1}=s]},
\]
so the LP directly identifies the causal impulse response without requiring specification of the full data-generating process. The paper further states that this causal interpretation is robust to the choice of state variable, whereas commonly used linear interaction LPs generally fail to recover causal objects [2605.05404].

Clustered LPs occupy an intermediate position. If the driving variables are exogenous, the clustered coefficient satisfies
\[
\beta_k^h = CAR^h(1,k)=CMR^h(k),
\]
where \(CAR^h(\delta,k)\) is the conditional average response in cluster \(k\). If the driving variables are endogenous, \(\beta_k^h\) remains interpretable as
\[
\beta_k^h=\int \omega_k(e)\,\Psi_k^{h\prime}(e)\,de,
\]
a weighted average of marginal effects with a nonnegative, hump-shaped weight function \(\omega_k\) [2604.18778].

This literature implies that “state dependence” is not a single estimand. Depending on the maintained assumptions, a state-dependent LP may identify a causal state-specific impulse response, a best linear approximation to a nonlinear response surface, or a weighted average of marginal effects.

## 3. Semiparametric and nonparametric estimation

The semiparametric approach of "Semiparametric Local Projections" defines the state-specific conditional mean \(g_{0,h}(x,z,s)\) nonparametrically and identifies the impulse response through a doubly robust moment. Let \(f_0(x\mid z,s)\) be the conditional density of \(x_t\) given \((z_{t-1},S_{t-1})\), and define the density ratio
\[
\alpha_0(x,z,s)
=
\frac{f_0(x-\delta\mid z,s)-f_0(x\mid z,s)}{f_0(x\mid z,s)}.
\]
Then, for \(\theta_h(s)\equiv IRF_h(s)\),
\[
E\Bigl[
g(x_t+\delta,z_{t-1},s)-g(x_t,z_{t-1},s)-\theta_h(s)
+\alpha(x_t,z_{t-1},s)\bigl(y_{t+h}-g(x_t,z_{t-1},s)\bigr)
\;\Bigm|\; S_{t-1}=s
\Bigr]=0.
\]
If either \(g=g_{0,h}\) or \(\alpha=\alpha_0\), the moment still identifies \(\theta_h(s)\); this is the double-robustness property [2606.13519].

Estimation proceeds in two stages. First, \(g_{0,h}(x,z,s)\) is estimated nonparametrically, for example by series or kernel methods over observations with \(S_{t-1}=s\). Second, \(\alpha_0(x,z,s)\) is estimated either by fitting conditional densities and forming
\[
\hat\alpha
=
\frac{\hat f(x-\delta\mid z,s)-\hat f(x\mid z,s)}{\hat f(x\mid z,s)},
\]
or by directly solving a minimum-distance problem using a rich dictionary \(\{b_j(x,z)\}\) and LASSO, as in Chernozhukov et al. (2022) [2606.13519].

To handle serial dependence, the paper uses cross-fitting for time series under an NLO scheme. The sample is split into \(K\) blocks of consecutive observations. For each block \(I_\ell\), nuisance functions are estimated on the “quasi-complement” excluding block \(\ell\) and its two immediate neighbors; the resulting blockwise moment
\[
\psi_{t,\ell}
=
\hat g_\ell(x_t+\delta,z_{t-1},S_{t-1})
-
\hat g_\ell(x_t,z_{t-1},S_{t-1})
+
\hat\alpha_\ell(x_t,z_{t-1},S_{t-1})
\bigl[y_{t+h}-\hat g_\ell(x_t,z_{t-1},S_{t-1})\bigr]
\]
is averaged over state-specific observations and then across folds. Under stationarity, geometric \(\beta\)-mixing, moment boundedness, and nuisance-rate conditions, the resulting estimator is \(\sqrt{T}\)-consistent and asymptotically normal, with consistent variance estimation via HAC methods such as a Bartlett kernel with Andrews’ bandwidth [2606.13519].

A separate nonparametric route is the sieve LP for micro-macro panels. Let \(\{\Phi_{1,J}(s),\dots,\Phi_{J,J}(s)\}\) be a sieve basis, such as cubic B-splines with \(J\) knots, and approximate
\[
g_h(s)\approx \sum_{j=1}^J b_{h,j}\Phi_{j,J}(s).
\]
The LP becomes
\[
Y_{i,t+h}
=
\sum_{j=1}^J b_{h,j}\bigl[\Phi_{j,J}(S_{i,t-1})\,\varepsilon_t\bigr]
+
W_{i,t-1}'\gamma_h
+
u_{i,t+h},
\]
estimated by OLS. The basis dimension can be selected by
\[
\mathrm{AIC}_h(J)
=
n\log\bigl(\hat\sigma_h^2(J)\bigr)+2\,(J+\dim W),
\qquad n=N(T-h),
\]
and under mixing, finite moments, and sieve approximation rate \(O(J^{-\kappa})\) with \(\kappa>1/2\), the paper establishes both a pointwise CLT and asymptotically valid uniform confidence bands [2605.05404].

These two approaches address different empirical environments. The semiparametric estimator targets nonlinear time-series responses with double robustness and explicit serial-dependence handling; the sieve LP targets causal state-varying slopes in micro-macro panels while allowing valid pointwise and uniform inference.

## 4. Approximation-based specifications and regime partitioning

A large empirical literature uses lower-dimensional approximations rather than fully nonparametric \(g_h(\cdot)\). You evaluates three prominent specifications in a quadratic-VAR laboratory. The linear benchmark is
\[
y_{t+h}=\alpha_h+\beta_h\varepsilon_t+u_{t+h}.
\]
A shock-sign interaction or asymmetric LP augments it to
\[
y_{t+h}
=
\alpha_h+\beta_h\varepsilon_t+\gamma_h I(\varepsilon_t>0)\varepsilon_t
+W_t'\pi_h+u_{t+h},
\]
with implied response
\[
\mathrm{IRF}^{\rm Asym}_h(\delta)=
\begin{cases}
(\beta_h+\gamma_h)\delta,& \delta>0,\\
\beta_h\delta,& \delta\le 0.
\end{cases}
\]
A lagged-state interaction uses
\[
y_{t+h}
=
\alpha_h+\beta_h^{(0)}\varepsilon_t+\beta_h^{(1)}s_{t-1}\varepsilon_t+W_t'\pi_h+u_{t+h},
\]
with
\[
\mathrm{IRF}^{\rm Lag}_h(s,\delta)=\bigl(\beta_h^{(0)}+\beta_h^{(1)}s\bigr)\delta.
\]
The augmented feasible specification combines both margins of nonlinearity:
\[
y_{t+h}
=
\alpha_h+\beta_h\varepsilon_t+\theta_h\varepsilon_t^2+\delta_h s_{t-1}\varepsilon_t+W_t'\pi_h+u_{t+h},
\]
so that
\[
\mathrm{IRF}^{\rm Feas}_h(s,\delta)
=
\beta_h\delta+\delta_h s\delta+\theta_h\delta^2.
\]
In the QVAR environment studied by You, this feasible specification exactly matches the true conditional average response [2602.14455].

Clustered LPs replace functional approximations by a data-driven partition of the state space. For a candidate number of clusters \(K\), the state observations \(\{s_t\}\) are classified by \(k\)-means:
\[
\min_{C_1,\dots,C_K,\mu_1,\dots,\mu_K}
\sum_{k=1}^K\sum_{t:s_t\in C_k}\|s_t-\mu_k\|^2.
\]
The resulting indicators \(D_{k,t}\) enter the LP, and impulse responses across horizons are estimated by GMM using stacked moment conditions
\[
E\bigl[\mathbf x_t(\mathbf y_t-\mathbf C\,\mathbf x_t)'\bigr]=0.
\]
The number of clusters is chosen iteratively: starting from a large \(K_{\max}\), LPs are estimated, clusters with indistinguishable impulse responses up to horizon \(\tilde H\) are tested by pairwise Wald tests with Bonferroni correction, and the closest pair is merged until all remaining clusters are mutually distinct [2604.18778].

These approximation-based designs are not equivalent. The asymmetric LP is directed at higher-order effects, the lag-interaction LP at dependence on observable proxies for latent states, the feasible augmented LP at both margins jointly, and the clustered LP at piecewise-constant heterogeneity over a low-dimensional state space.

## 5. VAR comparisons, IV complications, and inferential cautions

Winkler shows that state-dependent LPs and state-dependent VARs generally target different estimands. For a reduced-form VAR with state-dependent transition matrices, three objects can be defined: a fixed-state IRF,
\[
\theta^f_{\mathit{VAR},h}(s):=[\Pi(s)^hA(s)]_{21},
\]
a moving-state IRF,
\[
\theta^m_{\mathit{VAR},h}(s)
:=
\bigl[E(\Pi(S_{t+h-1})\cdots\Pi(S_t)\mid S_{t-1}=s)A(s)\bigr]_{21},
\]
and the LP estimand,
\[
\frac{E[Y_{t+h}X_t\mid S_{t-1}=s]}{E[X_t^2\mid S_{t-1}=s]}.
\]
In general,
\[
\theta^f_{\mathit{VAR},h}\neq \theta^m_{\mathit{VAR},h}\neq
\frac{E[Y_{t+h}X_t\mid S_{t-1}=s]}{E[X_t^2\mid S_{t-1}=s]},
\]
even with exogenous state and independent errors. To recover the LP estimand with VAR machinery, the paper proposes a back-shifted VAR estimator based on \(h+1\) separate VARs, each conditioning on a shifted lag of the state, yielding
\[
\theta^b_{\mathit{VAR},h}(s)
=
[\Pi^h(s)\Pi^{h-1}(s)\cdots\Pi^1(s)A^0(s)]_{21}
=
\frac{E[Y_{t+h}X_t\mid S_{t-1}=s]}{E[X_t^2\mid S_{t-1}=s]}.
\]
This matches the LP probability limit exactly [2601.01622].

The IV case introduces an additional source of misinterpretation. If
\[
Y_{t+h}=f(S_{t-1})X_t\beta^h+\text{error}
\]
is estimated by 2SLS using instruments \(f(S_{t-1})Z_t\), then the estimand depends not only on state variation in the structural response but also on state variation in the first stage. Winkler shows that state-dependent weighting can generate nonzero interaction terms even when the effects are not state-dependent. The familiar interpretation \(\beta_1^h=b(1)-b(0)\) requires either \(b(s)\) constant in \(s\) or a linear first stage \(X_t=\Theta(V_t)+Z_t\), so that \(X'(z,v)\) is state-invariant [2601.01622].

Inferential practice also changes under nonlinear specifications. In the feasible augmented LP studied by You, HAC/HAR standard errors are used because the squared-shock regressor induces serial correlation in the score that EHW cannot handle. More generally, both the semiparametric estimator and clustered LP rely on HAC-robust long-run variance estimators, and the panel sieve LP uses a HAC estimator of the long-run variance of the partialled-out sieve scores [2602.14455].

The common caution is that interaction coefficients, VAR IRFs, and IV interaction terms are not interchangeable objects. Their equality requires specific conditions rather than generic appeal to “state dependence.”

## 6. Evidence from simulations and empirical applications

Simulation evidence emphasizes that specification choice matters. In You’s quadratic-VAR laboratory, linear LPs fail to recover any of the state dependence or higher-order effects when shocks are symmetrically distributed. Shock-sign interactions recover part of the higher-order term and reduce errors only in large-\(|\delta|\) tail-shock regions. Shock–lagged-state interactions recover part of the state-dependence coefficient provided the observable proxy tracks the latent state well, with gains concentrated in tail-state regions. The augmented feasible specification recovers both components simultaneously and achieves the smallest approximation error across the entire joint distribution of shocks and states [2602.14455].

Clustered LPs perform well when the underlying heterogeneity is well approximated by regime partitions. In the Monte Carlo design with \(T=2000\), \(M=10{,}000\) replications, \(K_{\max}=10\), \(\tilde H=5\), and \(\alpha=0.05\), the method accurately recovers the piecewise-constant approximation to the conditional average response. The iterative algorithm tends to be conservative in small samples and merges clusters whose IRFs are not statistically distinguishable [2604.18778].

The semiparametric estimator of "Semiparametric Local Projections" is evaluated across a range of nonlinear data-generating processes. Its defining result is that the estimator is \(\sqrt{T}\)-consistent and asymptotically normal while remaining robust to first-stage nonparametric errors and serial dependence. The paper also reports two empirical examples, although the data block does not detail their substantive findings [2606.13519].

Three empirical applications illustrate how substantive conclusions can change once state dependence is modeled more flexibly. In the monetary-policy application of the clustered LP paper, the state vector combines macroeconomic uncertainty (\(\mathrm{MacroUcer}\)) and monetary policy uncertainty (\(\mathrm{MPU}\)); the iterative procedure selects four clusters: Low MacroUcer / Low MPU, Low MacroUcer / High MPU, Moderate MacroUcer / High MPU, and High MacroUcer / Low MPU. The estimates suggest that macroeconomic uncertainty primarily amplifies the risk compensation embedded in the term premium, while monetary policy uncertainty governs the speed and persistence with which markets revise expected future rates after a contractionary monetary policy shock [2604.18778].

In You’s application to Romer–Romer monetary shocks, the feasible augmented LP implies larger real effects in trough states than in peak states, while the quadratic term is often statistically significant but economically small. In David et al., a linear-interaction LP implies a monotone increase in firm investment responsiveness with distance-to-default, whereas the sieve LP reveals a hump-shaped response: very-distressed firms respond little, mid-range firms respond the most, and very safe firms respond less than mid-range firms. Because most firms sit near the mean, the hump-shaped nonparametric response implies a larger aggregate investment response than the linear LP, and the linear LP understates the bottom-up aggregate effect by up to an order of magnitude [2602.14455][2605.05404].

The principal limitations are likewise specification-specific. Clustered LPs rely on a piecewise-constant approximation, the “true” number of regimes need not exist, and performance depends on partition choice and sample size. Lag-based LPs hinge on how well the chosen observable proxies the latent state. More generally, the extra exogeneity restriction \(X_t\perp S_{t-1}\) is substantive rather than automatic, and applied work must justify it if interaction coefficients are to receive the weighted-average interpretation established by Winkler [2604.18778][2601.01622].

State-dependent LPs therefore form a family rather than a single estimator. Their modern theory distinguishes projection-based weighted averages, structurally causal state-specific responses, and approximation devices tailored to particular nonlinearities. The practical consequence is that empirical interpretation must be aligned with the exact specification, the maintained identifying assumptions, and the mode of heterogeneity being approximated.

Source: https://www.emergentmind.com/topics/state-dependent-local-projections-lps