---
title: Cumulative Cross-World Weighted Effects
url: https://www.emergentmind.com/topics/cumulative-cross-world-weighted-effects
type: topic
---

# Cumulative Cross-World Weighted Effects

Searching arXiv for the cited papers and closely related longitudinal causal inference work.
Cumulative cross-world weighted effects are longitudinal causal estimands for comparing two static treatment regimes when standard longitudinal positivity fails. They are defined as weighted averages of the potential outcome contrast $Y(\overline a_T)-Y(\overline a_T^\prime)$, with weights that accumulate over time and depend simultaneously on the treatment process under both counterfactual treatment histories [2507.10774]. The construction is intended to preserve the contrast between the two regimes themselves rather than substituting an implementable stochastic policy, but this comes at the cost of targeting a non-implementable cross-world quantity [2507.10774]. The topic also intersects with recent work on clone-censor-weighting (CCW), which shows that weighting procedures for treatment initiation windows can implicitly target hybrid or impossible interventions assembled from incompatible treatment histories, thereby clarifying why cross-world weighting is both interpretable in mechanistic terms and problematic in policy terms [2404.15073].

## 1. Formal longitudinal setup and estimand

The proposed framework is longitudinal with $T$ time points and observed data
\[
Z=(X_1,A_1,X_2,A_2,\dots,X_T,A_T,Y),
\]
where $X_t$ are time-varying covariates, $A_t\in\{0,1\}$ are binary treatments, and $Y$ is the final outcome [2507.10774]. Treatment histories are denoted $\overline A_t=(A_1,\dots,A_t)$, covariate histories are $\overline X_t=(X_1,\dots,X_t)$, and the history before treatment at time $t$ is
\[
H_t=(\overline X_t,\overline A_{t-1})
\]
[2507.10774].

Potential outcomes are defined in an NPSEM:
\[
X_t=f_{X,t}(A_{t-1},H_{t-1},U_{X,t}),\qquad A_t=f_{A,t}(H_t,U_{A,t}),\qquad Y=f_Y(A_T,H_T,U_Y)
\]
[2507.10774]. For a regime $\overline a_T=(a_1,\dots,a_T)$, the counterfactual outcome is $Y(\overline a_T)$, and the counterfactual covariate history is $\overline X_t(\overline a_{t-1})$ [2507.10774].

The cumulative cross-world weighted effect is
\[
\psi(\overline a_T,\overline a_T^\prime) := E\!\left( \{Y(\overline a_T)-Y(\overline a_T^\prime)\} \prod_{t=1}^T w_t\!\left[p_t\{\overline X_t(\overline a_{t-1})\}\right] w_t^\prime\!\left[p_t^\prime\{\overline X_t(\overline a_{t-1}^\prime)\}\right] \right)
\]
[2507.10774]. Here the natural propensity scores are
\[
p_t\{\overline X_t(\overline a_{t-1})\} = P\!\left\{A_t(\overline a_{t-1})=a_t\mid \overline X_t(\overline a_{t-1})\right\},
\]
\[
p_t^\prime\{\overline X_t(\overline a_{t-1}^\prime)\} = P\!\left\{A_t(\overline a_{t-1}^\prime)=a_t^\prime\mid \overline X_t(\overline a_{t-1}^\prime)\right\}
\]
[2507.10774]. Thus the estimand is not merely a causal contrast between two regime-specific means; it is a weighted contrast whose weighting functional itself is cross-world.

A plausible implication is that the estimand should be understood as selecting, across time, the portions of the two counterfactual trajectories that remain jointly informative when positivity fails. The paper states this point more directly by emphasizing that the weights adapt “cumulatively across timepoints and simultaneously across both counterfactual treatment histories” [2507.10774].

## 2. Cumulative weighting across counterfactual worlds

The “cumulative” feature is the product structure
\[
\prod_{t=1}^T w_t\!\left[p_t\{\overline X_t(\overline a_{t-1})\}\right] w_t^\prime\!\left[p_t^\prime\{\overline X_t(\overline a_{t-1}^\prime)\}\right]
\]
[2507.10774]. Weighting therefore responds at every time point to longitudinal nonpositivity rather than only at baseline.

The paper gives several explicit weighting choices [2507.10774]:

| Choice | $w_t$ | $w_t^\prime$ |
|---|---|---|
| No weighting | $1$ | $1$ |
| Weight toward only the target regime | $p_t$ | $1$ |
| Weight toward only the comparator regime | $1$ | $p_t^\prime$ |
| Overlap weighting | $p_t$ | $p_t^\prime$ |
| Trimming | $\mathbb{1}(p_t\ge \varepsilon)$ | $\mathbb{1}(p_t^\prime\ge \varepsilon^\prime)$ |

A concrete trimming example is
\[
w_t[p_t]=\mathbb{1}(p_t>\varepsilon),\qquad w_t^\prime[p_t^\prime]=\mathbb{1}(p_t^\prime>\varepsilon^\prime)
\]
[2507.10774]. A smooth trimming choice used in the empirical application is
\[
w_t(\pi_t)=1-e^{-20\pi_t},\qquad w_t^\prime(\pi_t^\prime)=1-e^{-20\pi_t^\prime}
\]
[2507.10774].

The key restriction on the weights is
\[
\pi_t(\overline x_t)\pi_t^\prime(\overline x_t)=0 \implies w_t\{\pi_t(\overline x_t)\}w_t^\prime\{\pi_t^\prime(\overline x_t)\}=0
\]
[2507.10774]. This requirement ensures that if either relevant propensity score is zero, the combined weight is also zero. In substantive terms, paths unsupported in either counterfactual world are removed from the weighted functional.

This suggests that weighting is not simply a variance-control device. It determines which regions of longitudinal state space contribute to the estimand and therefore partly defines the target itself.

## 3. Identification without standard positivity

Identification is developed under NPSEM and strong sequential randomization:
\[
U_{A,t}\ \perp\ \{\underline U_{X,t+1},\underline U_{A,t+1},U_Y\}\mid H_t, \qquad t=1,\dots,T
\]
[2507.10774]. The paper notes that this is stronger than ordinary sequential exchangeability because the estimand depends on cross-world quantities [2507.10774].

Under NPSEM and strong sequential randomization, if positivity has held up to earlier times, then the natural propensity scores are identified by observed-data quantities:
\[
P\{A_t(\overline a_{t-1})=a_t\mid \overline X_t(\overline a_{t-1})\} = P(A_t=a_t\mid \overline X_t,\overline A_{t-1}=\overline a_{t-1}) =:\pi_t(\overline X_t),
\]
with an analogous expression for $\pi_t^\prime(\overline X_t)$ under $\overline a_T^\prime$ [2507.10774].

The main identification formula is
\[
\psi(\overline a_T,\overline a_T^\prime) = \int E(Y\mid \overline a_T,\overline x_T) \prod_{t=1}^T w_t\{\pi_t(\overline x_t)\}w_t^\prime\{\pi_t^\prime(\overline x_t)\} \,dP(x_t\mid \overline A_{t-1}=\overline a_{t-1},\overline x_{t-1})
\]
\[
- \int E(Y\mid \overline a_T^\prime,\overline x_T) \prod_{t=1}^T w_t\{\pi_t(\overline x_t)\}w_t^\prime\{\pi_t^\prime(\overline x_t)\} \,dP(x_t\mid \overline A_{t-1}=\overline a_{t-1}^\prime,\overline x_{t-1})
\]
[2507.10774]. The paper explicitly states that no standard positivity assumption is required for identification because the design of the weights replaces it [2507.10774].

The framework distinguishes positivity from overlap of covariate distributions through the density ratio
\[
\rho_t(\overline X_t) = \frac{dP(X_t\mid \overline A_{t-1}=\overline a_{t-1},\overline X_{t-1})}
{dP(X_t\mid \overline A_{t-1}=\overline a_{t-1}^\prime,\overline X_{t-1})}
\]
and the condition
\[
P\{\rho_t(\overline X_t)\in (0,\infty)\}>0,\qquad \text{for every }t
\]
which the paper terms partial common support [2507.10774]. Partial common support is weaker than full overlap, but the paper states that it ensures the weighted functional remains informative about the causal contrast [2507.10774].

## 4. Interpretation, non-implementability, and the cross-world problem

The paper presents cumulative cross-world weighted effects as mechanism-relevant but not policy-relevant in the sense of a feasible intervention [2507.10774]. The reason is structural: the estimand depends on a product of weights evaluated under two counterfactual worlds simultaneously, and one cannot implement an intervention that generates both $\overline X_t(\overline a_{t-1})$ and $\overline X_t(\overline a_{t-1}^\prime)$ for the same person [2507.10774].

The resulting tradeoff is explicit. If positivity holds, the ordinary causal contrast
\[
E\{Y(\overline a_T)-Y(\overline a_T^\prime)\}
\]
is both interpretable and implementable. If positivity fails, implementable stochastic or modified policies can be identified, but those may conflate regime effects with downstream effects on intermediate treatments and covariates [2507.10774]. The cumulative cross-world weighted effect instead preserves the pure contrast of the two static regimes, but only as a non-implementable cross-world quantity [2507.10774].

The paper also gives a null-preservation property:
\[
P\{Y(\overline a_T)=Y(\overline a_T^\prime)\}=1 \quad\Rightarrow\quad \psi(\overline a_T,\overline a_T^\prime)=0
\]
[2507.10774]. This establishes that the estimand respects a strong form of causal nullity despite its non-implementable character.

Closely related concerns appear in the analysis of CCW for treatment initiation windows. There, a regimen such as “start treatment prior to day 30” is shown to estimate the potential outcome under a two-stage intervention where “A) prior to day 30, everyone follows the treatment start distribution of the study population and B) everyone who has not initiated by day 30 is forced to initiate on day 30” [2404.15073]. When earlier initiators are used to represent those forced to initiate at the end of the window, the target may require assigning a day-30 starter the prior exposure history of someone who started earlier, producing a cross-world or impossible intervention [2404.15073]. This parallel is conceptually important: in both settings, weighting may preserve a mechanistic contrast while moving away from interventions that could be literally enacted.

## 5. Partial common support and collapse to zero

A central substantive limitation is that identification alone does not guarantee informativeness. The paper states that if for some $t$,
\[
P\{\rho_t(\overline X_t)\in(0,\infty)\}=0,
\]
then
\[
\psi(\overline a_T,\overline a_T^\prime)=0
\]
regardless of the true potential outcome contrast [2507.10774]. The argument is that when the relevant conditioning event has probability zero, the corresponding propensity score is set to zero, so the combined weight is zero on those paths; if no common support remains, the entire weighted functional vanishes [2507.10774].

Conversely, under positive-probability overlap at each timepoint, the paper states that the functional can preserve sign: if
\[
P\{Y(\overline a_T)>Y(\overline a_T^\prime)\}=1,
\]
then
\[
\psi(\overline a_T,\overline a_T^\prime)>0,
\]
and if
\[
P\{Y(\overline a_T)<Y(\overline a_T^\prime)\}=1,
\]
then
\[
\psi(\overline a_T,\overline a_T^\prime)<0
\]
[2507.10774]. Partial common support is therefore not required for algebraic identification, but it is required for substantive interpretability [2507.10774].

This suggests that cumulative cross-world weighted effects answer a restricted version of the original causal question: they remain informative only on the longitudinal region where both counterfactual trajectories retain some common support. Outside that region, the estimand is intentionally silent.

## 6. Estimation theory and machine-learning implementation

For the identified target component $\psi(\overline a_T)$, the paper defines the backward recursive regression
\[
m_{T+1}=Y,
\]
\[
m_t(\overline X_t) = E\!\left[ m_{t+1}(\overline X_{t+1}) w_{t+1}\{\pi_{t+1}(\overline X_{t+1})\} w_{t+1}^\prime\{\pi_{t+1}^\prime(\overline X_{t+1})\} \mid \overline A_t=\overline a_t,\overline X_t \right]
\]
[2507.10774]. It also defines inverse weights
\[
r_t(A_t,\overline X_t) = \frac{\mathbb{1}(A_t=a_t)\,w_t\{\pi_t(\overline X_t)\}w_t^\prime\{\pi_t^\prime(\overline X_t)\}}{\pi_t(\overline X_t)},
\]
\[
r_t^\prime(A_t,\overline X_t) = \frac{\mathbb{1}(A_t=a_t^\prime)\,w_t\{\pi_t(\overline X_t)\}w_t^\prime\{\pi_t^\prime(\overline X_t)\}}{\pi_t^\prime(\overline X_t)}
\]
[2507.10774].

Under smooth weights and boundedness and moment assumptions, the uncentered efficient influence function is
\[
\varphi(Z)=\varphi_m(Z)+\varphi_w(Z)
\]
[2507.10774]. The paper notes that $\varphi_m$ is the usual sequential-regression residual term and $\varphi_w$ captures uncertainty from estimating the weights and propensity scores; the EIF includes terms involving both $r_t$ and $r_t^\prime$, plus derivative terms
\[
\phi_t(A_t,\overline X_t) = \dot w_t\{\pi_t(\overline X_t)\}\{1(A_t=a_t)-\pi_t(\overline X_t)\},
\]
\[
\phi_t^\prime(A_t,\overline X_t) = \dot w_t^\prime\{\pi_t^\prime(\overline X_t)\}\{1(A_t=a_t^\prime)-\pi_t^\prime(\overline X_t)\}
\]
[2507.10774]. A notable feature is that the EIF contains a term propagating information from the non-target regime through the density ratio $\rho_t$ [2507.10774].

The doubly robust-style algorithm uses: propensity estimation for both regimes, density-ratio estimation for $\rho_t$, backward regression for $m_t$, plug-in EIF evaluation on held-out data, and cross-fitting or sample splitting [2507.10774]. The estimator is
\[
\widehat\psi(\overline a_T)=P_n\{\widehat\varphi(Z)\}, \qquad \widehat\sigma^2=P_n\big[(\widehat\varphi-\widehat\psi)^2\big]
\]
[2507.10774]. Under the stated conditions,
\[
\sqrt{n}\,\frac{\widehat\psi(\overline a_T)-\psi(\overline a_T)}{\widehat\sigma} \rightsquigarrow N(0,1)
\]
[2507.10774]. The nuisance estimators need only converge at roughly $n^{-1/4}$ rate because the bias is second order and the remainder is $o_P(n^{-1/2})$ [2507.10774].

To avoid direct density-ratio estimation, the paper uses the identity
\[
\rho_t(\overline X_t) = \frac{P(\overline A_{t-1}=\overline a_{t-1}\mid \overline X_t)\,
P(\overline A_{t-1}^\prime=\overline a_{t-1}^\prime\mid \overline X_{t-1})}
{P(\overline A_{t-1}=\overline a_{t-1}^\prime\mid \overline X_t)\,
P(\overline A_{t-1}=\overline a_{t-1}\mid \overline X_{t-1})}
\]
so that $\rho_t$ can be estimated through standard binary regressions for conditional probabilities [2507.10774]. The paper explicitly characterizes this reformulation as machine-learning friendly [2507.10774].

## 7. Relation to treatment-window CCW and empirical illustration

The connection to CCW clarifies a broader family resemblance among longitudinal weighting estimands. In CCW, each eligible individual is cloned into the regimens under study, clones are censored when they become incompatible with the regimen, and inverse probability of censoring weights reweight the remaining uncensored person-time [2404.15073]. For a “start by day 30” regimen, those who have not initiated by day 30 are censored at day 30, and the weights are based on the inverse of the conditional probability of remaining uncensored:
\[
W_i(t)=\prod_{k \le t}\frac{1}{P\{C_i(k)=0 \mid \bar C_i(k-1)=0, \text{covariate and exposure history}\}}
\]
[2404.15073].

The paper on CCW emphasizes that when exposure effects are time-varying, ignoring exposure history when estimating IPCW can estimate the risk under an impossible intervention and can create selection bias [2404.15073]. Its “limited CCW” versus “all initiator CCW” distinction formalizes this point: the all-initiator approach can estimate the risk under an impossible intervention when exposure effects vary over time because it retroactively gives day-30 starters the prior exposure history of earlier starters [2404.15073]. The simplifying assumptions under which this problem disappears are “No exposure effect,” “Exposure effect begins after the end of the period,” and “Instantaneous effect,” with different implications for whether earlier initiators may be included in IPCW and whether the distribution of treatment timing can be ignored [2404.15073].

A plausible implication is that cumulative cross-world weighted effects and CCW illuminate the same conceptual boundary from different directions. CCW shows how cross-world structure can arise inadvertently from weighting within treatment initiation windows; cumulative cross-world weighted effects make that structure explicit and treat it as the target of inference.

The empirical illustration in the cumulative cross-world weighting paper uses the wagepan dataset, following Vella and Verbeek, with worker data from 1980–1987 and analysis focused on 1980–1983 [2507.10774]. The treatment is union membership each year, the target regimes are always unionized $(1,1,1,1)$ and never unionized $(0,0,0,0)$, and the outcome is log wage in 1983 [2507.10774]. Estimation uses smooth trimming,
\[
w_t(\pi_t)=1-\exp(-20\pi_t),\qquad
w_t^\prime(\pi_t^\prime)=1-\exp(-20\pi_t^\prime),
\]
together with a doubly robust-style estimator, five-fold cross-fitting, and SuperLearner nuisance estimation with linear model, GLM, lasso, regression tree, and random forest learners [2507.10774].

The reported effect estimate is
\[
\widehat\psi = 0.216,
\]
with 95% CI
\[
[0.046,\ 0.387]
\]
[2507.10774]. On the log-wage scale, the paper states that this corresponds to roughly a 22% increase in wages in 1983 attributable to always being in a union versus never being in a union, under the paper’s assumptions [2507.10774]. The authors note, however, that the no-unmeasured-confounding assumption may be questionable and that sensitivity analysis would be valuable [2507.10774].

Cumulative cross-world weighted effects therefore occupy a specific niche in longitudinal causal inference: they are cumulative products of time-specific weights, cross-world because they use propensity information from two counterfactual treatment histories, identifiable without standard positivity, mechanistically faithful but not implementable, and substantively informative only under partial common support [2507.10774].

Source: https://www.emergentmind.com/topics/cumulative-cross-world-weighted-effects