Stochastic Incremental Propensity Score Interventions
- The paper develops a framework that shifts treatment odds via a user-specified factor δ, enabling realistic causal effect estimation without forcing deterministic treatment.
- It utilizes semiparametric efficiency theory and cross-fitting techniques to achieve robust and precise inference in both single- and longitudinal treatment settings.
- The methodology extends to continuous exposures, policy learning, and sensitivity analysis, offering practical improvements in handling positivity violations and treatment heterogeneity.
Stochastic incremental propensity score interventions are a class of stochastic interventions designed to quantify the effect of making treatment more or less likely, rather than forcing everyone deterministically into treatment or control. In the canonical binary-treatment case, they replace the observational propensity score with a shifted propensity
so that treatment odds are multiplied by a user-chosen factor . This produces a family of causal effects indexed by intervention strength, avoids the usual positivity bottleneck for many identification problems, and has been developed for single-time-point, longitudinal, conditional, continuous-exposure, and policy-learning settings (Kennedy, 2017, Bonvini et al., 2021, Zhao et al., 2023).
1. Definition and conceptual basis
The defining feature of a stochastic incremental propensity score intervention is that it does not set treatment to a fixed value. Instead, it modifies the observed treatment mechanism. For a single binary treatment , the intervention is specified by multiplying the observational odds of treatment by : Equivalently,
Thus increases the odds of treatment, decreases them, and reproduces the observed data-generating treatment mechanism (Bonvini et al., 2021, Kennedy, 2017).
This formulation differs fundamentally from deterministic interventions such as “treat everyone” or “treat no one.” Deterministic contrasts target outcomes under extreme treatment assignments, whereas incremental interventions ask what would happen if everyone’s odds of treatment were shifted by a factor 0. Several papers emphasize that this can be more realistic in applied settings where policy changes alter the likelihood of treatment rather than imposing universal compliance (Bonvini et al., 2021, Jacobs et al., 2023).
The same idea extends naturally to dynamic settings. In longitudinal data with treatment 1 at time 2 and history 3, the observed propensity 4 is replaced by
5
The intervention remains dynamic because it conditions on the realized history 6, but it is incremental because it shifts treatment propensity smoothly rather than assigning a fixed action (Kennedy, 2017, Kim et al., 2019).
This stochastic construction has an important interpretive consequence. The estimand is a curve 7, not a single contrast. In longitudinal studies this “single curve” replaces the need to index a potentially enormous collection of deterministic treatment regimes, including the 8 treatment histories that arise with binary treatment observed at 9 time points (Kennedy, 2017).
2. Causal estimands and identification without positivity
The standard target parameter for a single time point is the mean potential outcome under the shifted treatment mechanism: 0 Under consistency and exchangeability, this is identified by either a regression/g-computation-style formula or an inverse-probability-weighted formula: 1 where 2 (Bonvini et al., 2021). The positivity-free policy-learning formulation uses the same structure, writing the value of a stochastic policy 3 as
4
with the policy class restricted to those 5 induced by an incremental odds shift (Zhao et al., 2023).
The central identification property is that positivity is not required. If 6, then 7; if 8, then 9. The intervention therefore respects structural zeros in the observed treatment process and never creates treatment mass where none existed. This is the key reason these interventions remain identified under consistency and exchangeability or unconfoundedness alone (Bonvini et al., 2021, Zhao et al., 2023).
By contrast, ordinary deterministic policy-learning or treatment-effect formulas involve terms such as
0
which can blow up or become undefined when the denominator is zero. Incremental stochastic interventions avoid that failure because they stay within the support of the observed treatment mechanism (Zhao et al., 2023).
The relation to the average treatment effect is subtle. Incremental effects are not the ATE; they answer a different causal question. However, if positivity holds, then 1 as 2 and 3 as 4. In that sense, the ATE is nested as a limiting special case, while the incremental estimand remains stochastic and typically better behaved under practical overlap violations (Bonvini et al., 2021).
3. Semiparametric efficiency, estimation, and inference
A recurring theme in this literature is that incremental effects are smooth functionals of the observed law and admit efficient influence functions. For the single-time-point positivity-free policy-learning target
5
the efficient influence function is
6
with 7 (Zhao et al., 2023). More general efficient influence functions have been derived for longitudinal interventions, conditional incremental effects, continuous exposures, and sensitivity bounds (Kennedy, 2017, McClean et al., 2022, Schindl et al., 2024, Shen et al., 25 Jan 2026).
These influence functions motivate one-step and cross-fitted estimators. In the single-time-point policy-learning setting, the one-step estimator is
8
and the cross-fitted version is
9
where nuisance functions are estimated on training folds and evaluated on held-out folds (Zhao et al., 2023). The review literature likewise recommends sample splitting and cross-fitting to control overfitting bias and avoid Donsker-type restrictions (Bonvini et al., 2021).
The resulting estimators have second-order remainder terms. In the positivity-free policy-learning paper, if
0
then the one-step and cross-fitted estimators satisfy
1
and achieve the semiparametric efficiency bound (Zhao et al., 2023). The review chapter states the same qualitative conclusion for fixed 2: if 3 and 4 each converge faster than 5, then root-6 inference is still possible, with Wald-type confidence intervals and multiplier-bootstrap uniform confidence bands over a range of 7 values (Bonvini et al., 2021).
The literature does, however, distinguish between settings. Several single-time-point developments describe the estimators as doubly robust or orthogonal in spirit because the bias is of the order of products of nuisance estimation errors (Zhao et al., 2023, Bonvini et al., 2021). The original longitudinal curve paper emphasizes an important caveat: double robustness is not available in general because the estimand itself depends on the propensity score, so both the propensity score and the pseudo-regressions matter (Kennedy, 2017).
4. Variants and generalizations
One major extension is positivity-free policy learning with observational data. Instead of a deterministic rule 8, the policy is a stochastic treatment probability 9 defined by
0
The policy-learning objective is to find
1
possibly subject to constraints 2, with examples including fairness, budget/resource limits, and protection of vulnerable groups. Asymptotic guarantees are proved for parametric convex policy classes, where regret and plug-in value estimation are 3, and for Glivenko–Cantelli policy classes, where regret and value estimation are 4 (Zhao et al., 2023).
A second extension concerns conditional heterogeneity. The conditional incremental effect is
5
and the conditional incremental derivative effect is
6
The identified form of the derivative effect is
7
To estimate these effects, the literature develops a Projection-Learner, a flexible I-DR-Learner, and a heterogeneity summary based on the variance of the CIDE, denoted V-CIDE (McClean et al., 2022).
A third extension treats continuous exposures. Instead of multiplying binary treatment odds, the observed conditional density 8 is exponentially tilted: 9 This remains a smooth stochastic intervention because if 0 then 1 too. The mean outcome under the tilted law is
2
The paper derives the efficient influence function, shows that the semiparametric efficiency bound depends heavily on the squared likelihood ratio 3, proves a minimax lower bound of order 4, and interprets this as an effective sample size of 5 for large tilts (Schindl et al., 2024).
A fourth extension generalizes incremental interventions to discrete multi-arm treatments with cost-sensitive structure. “Interpolated stochastic interventions based on propensity scores, target policies and treatment-specific costs” defines a cost-penalized I-projection whose unique minimizer has the Boltzmann–Gibbs coupling
6
The induced source-tilted family 7 recovers standard incremental propensity score interventions in the binary Hamming-cost case, retains identification without global positivity for the source-tilted estimand, and yields one-step estimators with uniform confidence bands (Aguas, 14 Nov 2025).
Related work in stochastic intervention effect estimation uses the same odds-multiplicative shift
8
but frames it as a stochastic propensity score inside reinforcement-learning or genetic-search procedures for optimizing intervention strength 9 (Duong et al., 2021, Duong et al., 2021).
5. Longitudinal settings, dropout, and unmeasured confounding
Incremental interventions were developed in part to address the difficulties of longitudinal causal inference. With treatment at multiple time points, deterministic regimes can be non-identified under positivity violations and can be hard to estimate because the number of treatment histories grows exponentially in 0. Incremental interventions replace the full regime collection with a one-dimensional curve 1 while preserving a dynamic dependence on history (Kennedy, 2017).
In the longitudinal binary-treatment case, the target estimand is
2
with stage-specific shifted propensities
3
Under consistency and exchangeability, the mean counterfactual outcome is identified by a g-formula-type expression, and efficient influence functions can be derived using recursive pseudo-regressions 4. Cross-fitted influence-function estimators are 5-consistent and asymptotically normal uniformly over 6 under product-rate conditions on nuisance estimation errors (Kennedy, 2017).
The framework has also been extended to monotone dropout. In that setting, treatment positivity is still not required, but a positivity condition is imposed for dropout only: 7 The efficient influence function then combines incremental treatment weights with inverse dropout weights and recursive regressions. The resulting estimators remain uniformly asymptotically normal and can accommodate flexible nuisance estimation (Kim et al., 2019).
A particularly striking result in the many-timepoint setting is the variance ratio analysis. Under a simplified infinite-horizon regime, the variance ratio of incremental-effect estimators relative to deterministic always-treated or never-treated estimators decays exponentially in 8. The paper interprets this as near-exponential gains in statistical precision, because deterministic inverse-probability weights concentrate on rare treatment paths whereas incremental interventions spread mass across many paths (Kim et al., 2019).
Sensitivity analysis for unmeasured confounding has recently been developed for these effects. In the single-time-point setting, the framework adopts Rosenbaum’s sensitivity model, introduces conditional upper and lower bounds 9 for the unidentifiable components 0 and 1, and defines incremental-effect bounds
2
A cross-fitted doubly robust estimator is derived, and the bound estimators are asymptotically normal under mild conditions on nuisance estimation (Shen et al., 25 Jan 2026). The paper also shows that incremental effect bounds can be narrower or wider than those for mean potential outcomes, and that they must lie between the expected minimum and maximum of the conditional bounds on 3 and 4 (Shen et al., 25 Jan 2026).
For time-varying treatments, the same paper studies the marginal sensitivity model. Sharp bounds for incremental effects are identifiable from longitudinal data under this model, but a practical estimator has not yet been established. The stated obstacle is that the resulting optimization problem is difficult to dualize because one sensitivity function enters only implicitly through iterated expectations (Shen et al., 25 Jan 2026).
6. Applications, interpretation, and limitations
Applied work has used stochastic incremental propensity score interventions in criminology, health services research, sociology, and causal transport problems. In a criminology application to homelessness and rearrest among people on probation, the estimand is the average outcome if everyone’s odds of exposure were multiplied by some factor 5. With 6, reducing homelessness by roughly 65% corresponded to a 9% reduction in the estimated average rate of recidivism, while more moderate interventions were smaller and not statistically significant (Jacobs et al., 2023).
The original longitudinal application studied the effect of incarceration on marriage over 10 yearly time points in the National Longitudinal Survey of Youth 1997. Using random forests, 10-fold cross-fitting, and 10,000 bootstrap replications, the paper found strong evidence that incarceration reduces marriage probability; the null of no incremental effect was rejected with 7 over 8, and the estimated curve was nonlinear, with stronger effects for increasing incarceration than for decreasing it (Kennedy, 2017).
Other applications emphasize situations in which deterministic interventions are unrealistic or positivity is badly violated. The SPOTlight ICU admission study used conditional incremental effects because ICU capacity is limited and treatment assignment is subject to clinical judgment; estimated mortality was lowest near 9, suggesting the observed ICU admission process was near-optimal among incremental interventions (McClean et al., 2022). The Diabetes 130-US hospitals analysis used race as the sensitive attribute and demographic parity constraints; the reported finding was that standard methods, especially IPW, become highly unstable when positivity is severe, while the positivity-free methods were consistently better (Zhao et al., 2023). The low-dose aspirin study concluded that the estimated curves were fairly flat and confidence bands included no effect, illustrating that the intervention remains estimable despite severe noncompliance and dropout (Kim et al., 2019). The Add Health analysis reported robustness of the victimization–offending association up to about 0 in the main sensitivity analysis and around 1 at the 95% confidence level (Shen et al., 25 Jan 2026).
The framework has also been used indirectly to approximate treatment effects when treatment is absent in the target study. In that setting, a post-treatment exposure 2 is assigned according to a stochastic law meant to reproduce the treatment-induced distribution of 3. The data-integration paper stresses that this stochastic exposure effect equals the target treatment effect only under a strong “oracle exposure law” 4, and proposes instead to learn the treatment-induced exposure distribution from an external study and transport it to the target population (Wen et al., 26 Sep 2025).
Several limitations recur across the literature. First, incremental effects are not conventional deterministic intervention effects, so they answer a different causal question than the ATE (Bonvini et al., 2021). Second, interpretation hinges on the odds multiplier 5, which is intuitive but still abstract compared with a simple treated-versus-untreated contrast (Bonvini et al., 2021). Third, while positivity is not needed for identification in the core incremental setting, valid estimation still depends on accurately estimating nuisance functions such as propensity scores and outcome regressions, and precision can deteriorate when these are poorly estimated (Bonvini et al., 2021). Fourth, for continuous exposures, larger tilts make estimation intrinsically harder, with the optimal rate scaling as 6 and the effective sample size behaving like 7 (Schindl et al., 2024). Finally, the original longitudinal paper explicitly characterizes these effects as more descriptive than prescriptive, reflecting a tradeoff between robustness and the direct interpretability of hard intervention contrasts (Kennedy, 2017).
Taken together, these developments establish stochastic incremental propensity score interventions as a technically coherent family of stochastic causal estimands: they preserve the support of the observed treatment process, relax the need for treatment positivity in many settings, admit semiparametric efficiency theory and machine-learning-based estimation, and support extensions to policy learning, conditional heterogeneity, continuous exposures, dropout, sensitivity analysis, and cost-aware stochastic policy design (Kennedy, 2017, Zhao et al., 2023, Schindl et al., 2024, Aguas, 14 Nov 2025).