Papers
Topics
Authors
Recent
Search
2000 character limit reached

Decomposition of Spillover Effects Under Misspecification:Pseudo-true Estimands and a Local--Global Extension

Published 12 Feb 2026 in econ.EM, math.ST, and stat.ML | (2602.12023v1)

Abstract: Applied work with interference typically models outcomes as functions of own treatment and a low-dimensional exposure mapping of others' treatments, even when that mapping may be misspecified. This raises a basic question: what policy object are exposure-based estimands implicitly targeting, and how should we interpret their direct and spillover components relative to the underlying policy question? We take as primitive the marginal policy effect, defined as the effect of a small change in the treatment probability under the actual experimental design, and show that any researcher-chosen exposure mapping induces a unique pseudo-true outcome model. This model is the best approximation to the underlying potential outcomes that depends only on the user-chosen exposure. Utilizing that representation, the marginal policy effect admits a canonical decomposition into exposure-based direct and spillover effects, and each component provides its optimal approximation to the corresponding oracle objects that would be available if interference were fully known. We then focus on a setting that nests important empirical and theoretical applications in which both local network spillovers and global spillovers, such as market equilibrium, operate. There, the marginal policy effect further decomposes asymptotically into direct, local, and global channels. An important implication is that many existing methods are more robust than previously understood once we reinterpret their targets as channel-specific components of this pseudo-true policy estimand. Simulations and a semi-synthetic experiment calibrated to a large cash-transfer experiment show that these components can be recovered in realistic experimental designs.

Authors (2)

Summary

  • The paper shows that exposure-based analyses under misspecification target unique pseudo-true estimands, preserving the direct–indirect decomposition of marginal policy effects while quantifying approximation error through residual conditional variance.
  • The paper extends this framework to sparse-network and market-equilibrium settings, proving that marginal policy effects asymptotically separate into direct, local, and global spillovers because interactions between high-dimensional local exposure and low-dimensional prices are second order.
  • The paper establishes channel-specific robustness: Li–Wager network estimators consistently recover local spillovers despite unmodeled global interference, while augmented-IV estimators recover global spillovers despite unmodeled local interference, with rates determined by network sparsity and perturbation size.

Overview and motivation

This paper, by Yechan Park and Xiaodong Yang (2602.12023), addresses a foundational question in the econometrics of interference: when a researcher summarizes interference through an exposure mapping that is inevitably misspecified, what policy object do exposure-based estimands actually target? The authors take as primitive the marginal policy effect (MPE)—the welfare change from a small shift in treatment probability under the actual experimental design—and show that any researcher-chosen exposure mapping induces a unique "pseudo-true" outcome model, defined as the best mean-squared approximation to the true potential outcomes among all models depending on assignments only through the chosen exposure. Within this pseudo-true model, the MPE admits a canonical decomposition into direct and spillover components, each optimally approximating its oracle counterpart.

The paper's central claim is twofold. First, the Hu–Li–Wager identity—under which the marginal policy effect equals the sum of average direct and indirect effects under Bernoulli randomization—survives misspecification exactly once effects are reinterpreted as pseudo-true objects. Second, in structured environments with both local network spillovers and global equilibrium spillovers, the marginal policy effect decomposes asymptotically into three channels: direct, local, and global. A notable implication is that existing estimators are more robust than previously understood: Li–Wager-type network estimators remain consistent for the local component even in the presence of unmodeled global market interference, and Munro-type augmented-IV estimators remain consistent for the global component even with unmodeled local interference.

Pseudo-true estimands under misspecified exposures

The setup considers nn units with potential outcomes yi:{0,1}n→Ry_i:\{0,1\}^n\to\mathbb{R} of arbitrary complexity, assigned via a Bernoulli randomized controlled trial RCT(π)\mathrm{RCT}(\pi). The oracle estimands—the average direct effect (ADE) and average indirect effect (AIE) defined from the full potential outcome schedule—are generally intractable because they require knowledge of outcomes on exponentially many assignment vectors. Applied work instead posits an exposure mapping did_i and fits outcome models $h_i(d_i(\bw))$.

The key construction replaces yiy_i with the pseudo-true outcome

$\tilde{y}_i(\bw;\pi) = \mathbb{E}_{\bW^{(2)}\sim\mathrm{RCT}(\pi)}\bigl[y_i(\bW^{(2)}) \mid d_i(\bW^{(2)}) = d_i(\bw)\bigr],$

where $\bW^{(2)}$ is an independent second copy of the assignment vector. This is the unique minimizer of the design-based mean-squared discrepancy between yiy_i and any function of the chosen exposure. The construction parallels classical pseudo-true parameters in misspecified likelihood and GMM settings (White's KL projection; Hansen–Jagannathan distance).

Three results organize this section:

  1. Exact decomposition survives misspecification. Under RCT(π)\mathrm{RCT}(\pi), the pseudo-true MPE satisfies yi:{0,1}n→Ry_i:\{0,1\}^n\to\mathbb{R}0, extending the Hu–Li–Wager identity to arbitrary misspecified exposures. This means every exposure-based analysis implicitly targets a coherent decomposition of a well-defined policy derivative, not merely descriptive contrasts.
  2. Lipschitz continuity of estimands in the outcome model. For any candidate outcome functions yi:{0,1}n→Ry_i:\{0,1\}^n\to\mathbb{R}1, the induced functionals satisfy

yi:{0,1}n→Ry_i:\{0,1\}^n\to\mathbb{R}2

and the dependence on yi:{0,1}n→Ry_i:\{0,1\}^n\to\mathbb{R}3 cannot be improved (a linear-outcome counterexample attains the bound). Consequently, any method that approximates individual outcomes well also approximates the marginal policy effect well.

  1. Optimality of the pseudo-true model. Combining these,

yi:{0,1}n→Ry_i:\{0,1\}^n\to\mathbb{R}4

so if residual conditional variance is yi:{0,1}n→Ry_i:\{0,1\}^n\to\mathbb{R}5, the pseudo-true estimands converge to their oracle counterparts. The practical implication is that flexible nuisance estimators—IPW, regression adjustment, or modern machine learning for conditional expectations—can approximate yi:{0,1}n→Ry_i:\{0,1\}^n\to\mathbb{R}6 without modeling full interference structure.

The paper is careful to distinguish these estimands from generic exposure-contrast estimands comparing average outcomes at two exposure values: the pseudo-true objects are built from explicit perturbations of individual treatment statuses, so each contrast corresponds to a well-defined hypothetical intervention even under misspecification. The authors also concede, following Leung and Auerbach–Tabord-Meehan, that without additional structure one should not expect sharp identification of finer channels beyond what the exposure mapping encodes—a limitation that motivates the structured extension below.

The local–global environment

The second half specializes to a model class yi:{0,1}n→Ry_i:\{0,1\}^n\to\mathbb{R}7 nesting both empirical applications (cash transfers with price effects, informal insurance networks) and theoretical work on network and equilibrium interference. Outcomes take the form

yi:{0,1}n→Ry_i:\{0,1\}^n\to\mathbb{R}8

where yi:{0,1}n→Ry_i:\{0,1\}^n\to\mathbb{R}9 is the share of treated neighbors on a sparse graphon network and RCT(π)\mathrm{RCT}(\pi)0 is an equilibrium price vector solving aggregate excess demand RCT(π)\mathrm{RCT}(\pi)1. Units are drawn i.i.d. from a superpopulation; the graphon is sparse (RCT(π)\mathrm{RCT}(\pi)2 with RCT(π)\mathrm{RCT}(\pi)3, RCT(π)\mathrm{RCT}(\pi)4) and low-rank (rank RCT(π)\mathrm{RCT}(\pi)5); and the experimenter can augment the trial with individualized price perturbations RCT(π)\mathrm{RCT}(\pi)6, RCT(π)\mathrm{RCT}(\pi)7, RCT(π)\mathrm{RCT}(\pi)8, providing IV-like variation in the global state.

Main estimand result. As RCT(Ï€)\mathrm{RCT}(\pi)9, the finite-sample estimands converge to population limits:

Component Limit
Direct did_i0
Local spillover did_i1
Global spillover did_i2
Total (MPE) Sum of the three

Here did_i3 and did_i4 are population gradients of excess demand and outcomes at the clearing price did_i5. The striking feature—which the authors flag as surprising—is that the total effect decomposes additively into these three limits even though the finite-sample definition does not naturally split this way. The decoupling holds because the global channel operates through a low-dimensional consensus statistic fluctuating at order did_i6 while the local channel operates through high-dimensional ego exposures; their interaction is second order. Notably, the local limit coincides with Li–Wager's quantity at fixed price, and the global limit coincides with Munro et al.'s quantity at fixed local exposure—so each strand of the literature is, in fact, consistently estimating one channel-specific component of a single pseudo-true policy estimand.

A technical contribution worth noting: the proofs require a second-order expansion of the equilibrium price, did_i7, which strengthens Munro et al.'s market-clearing tolerance from did_i8 to did_i9. In doing so, the authors identify and repair a gap in the proof of Lemma 16 of Munro et al., where the quadratic Taylor remainder was controlled using an incorrect rate.

Estimators and rates

Three estimators correspond to the three components:

  • Direct effect: the Horvitz–Thompson estimator $h_i(d_i(\bw))$0, unbiased under RCT. Its limiting variance now includes contributions from both the local channel ($h_i(d_i(\bw))$1, a graphon-smoothed direct-effect gradient) and the global channel ($h_i(d_i(\bw))$2, a price-elasticity term), so standard errors must account for both mechanisms.
  • Local spillover: the PC-balancing estimator of Li–Wager, projecting the raw neighbor-treatment weights onto the subspace orthogonal to the top-$h_i(d_i(\bw))$3 eigenvectors of the adjacency matrix. It converges at rate $h_i(d_i(\bw))$4 around $h_i(d_i(\bw))$5 with the same variance as in the purely local setting—establishing robustness to unmodeled market interference.
  • Global spillover: an IV estimator combining price elasticities $h_i(d_i(\bw))$6 from the augmented perturbations with a Horvitz–Thompson estimate of the treatment effect on excess demands, giving $h_i(d_i(\bw))$7. It converges at rate $h_i(d_i(\bw))$8 around $h_i(d_i(\bw))$9, again robust to unmodeled local interference.

Summing the three yields a consistent estimator of the total MPE whose overall rate depends on whether yiy_i0: if yiy_i1 the local component dominates (rate yiy_i2); otherwise the global component dominates (rate yiy_i3). This trade-off makes explicit how experimental design choices—network density and perturbation magnitude—allocate power across channels.

Numerical evidence

Simulations use a fixed-index model with outcomes yiy_i4 across five link functions (linear, quadratic, cosine, logarithmic, cubic polynomial), varying the mixing parameter yiy_i5, treatment rate yiy_i6, and Erdős–Rényi density. With yiy_i7, Monte Carlo averages track the oracle ADE, local AIE, and global AIE closely across all link functions and parameter regimes. Log-log MSE plots over yiy_i8 confirm the predicted convergence rates in both sparse-network regimes considered.

The semi-synthetic application calibrates to the Philippine cash-transfer experiment of Filmer et al., extending Munro et al.'s egg-market calibration with a household network built from geographic blocks, homophily in housing and socioeconomic characteristics, and triadic closure (target density yiy_i9). With $\tilde{y}_i(\bw;\pi) = \mathbb{E}_{\bW^{(2)}\sim\mathrm{RCT}(\pi)}\bigl[y_i(\bW^{(2)}) \mid d_i(\bW^{(2)}) = d_i(\bw)\bigr],$0 households and 2,000 replications, the components are recovered accurately:

Estimator Truth Mean Bias SD
ADE 0.3514 0.3151 −0.0363 0.1522
AIE (local) −1.0871 −1.0732 0.0139 0.5731
AIE (global) −0.1333 −0.1320 0.0013 0.0973

In this calibration the local component is sizeable and negative—treated neighbors reduce the marginal gains from one's own transfer—and dominates the positive direct effect, so the calibrated total marginal policy effect is negative despite positive direct gains. The authors present this as one plausible configuration rather than an empirical finding, noting ex ante ambiguity in the sign of the total effect. Sensitivity checks over network densities $\tilde{y}_i(\bw;\pi) = \mathbb{E}_{\bW^{(2)}\sim\mathrm{RCT}(\pi)}\bigl[y_i(\bW^{(2)}) \mid d_i(\bW^{(2)}) = d_i(\bw)\bigr],$1 show ADE and global AIE stable while the local AIE varies with density, as expected given its dependence on local exposure.

Limitations and open questions

Several restrictions bound the scope of the results. The pseudo-true framework delivers only approximation guarantees relative to oracle estimands; closeness requires the chosen exposure to leave little residual conditional variance, and no test of this condition is provided. The local–global theory relies on specific structural assumptions: a low-rank sparse graphon with $\tilde{y}_i(\bw;\pi) = \mathbb{E}_{\bW^{(2)}\sim\mathrm{RCT}(\pi)}\bigl[y_i(\bW^{(2)}) \mid d_i(\bW^{(2)}) = d_i(\bw)\bigr],$2, smoothness bounds up to third derivatives, a unique population-clearing price with full-rank Jacobian, and availability of the augmented randomized trial with perturbation scale $\tilde{y}_i(\bw;\pi) = \mathbb{E}_{\bW^{(2)}\sim\mathrm{RCT}(\pi)}\bigl[y_i(\bW^{(2)}) \mid d_i(\bW^{(2)}) = d_i(\bw)\bigr],$3, $\tilde{y}_i(\bw;\pi) = \mathbb{E}_{\bW^{(2)}\sim\mathrm{RCT}(\pi)}\bigl[y_i(\bW^{(2)}) \mid d_i(\bW^{(2)}) = d_i(\bw)\bigr],$4—an experimental capability many field settings lack. The analysis covers only infinitesimal changes in the treatment rate; extension to other estimands such as the global average treatment effect remains open. Finally, the semi-synthetic calibration treats the estimated local coefficient $\tilde{y}_i(\bw;\pi) = \mathbb{E}_{\bW^{(2)}\sim\mathrm{RCT}(\pi)}\bigl[y_i(\bW^{(2)}) \mid d_i(\bW^{(2)}) = d_i(\bw)\bigr],$5 as fixed and depends on a single network draw, so it evaluates estimator performance under the assumed data-generating process rather than validating the structural model itself.

Conclusion

This paper provides a principled answer to what exposure-based spillover analyses estimate when their exposure mappings are misspecified: the canonical direct–indirect decomposition of the marginal policy effect within the best exposure-based approximation to true outcomes, with quantified error to oracle targets. In a structured local–global environment, it further shows the marginal policy effect splits asymptotically into direct, local, and global channels, and that leading estimators from the network-interference and market-equilibrium literatures remain valid for their respective components even when the other mechanism operates unmodeled. Simulations and a cash-transfer-calibrated experiment support recoverability in realistic designs. The framework reframes apparent fragility of existing methods as channel-specific consistency, and leaves open the design question of how to allocate experimental power—through saturation schemes, graph cluster randomization, or perturbation magnitudes—across the identified channels.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.