---
title: Runtime Confounding in Causal Models
url: https://www.emergentmind.com/topics/runtime-confounding
type: topic
---

# Runtime Confounding in Causal Models

Searching arXiv for recent and directly relevant papers on runtime confounding and adjacent formulations.
Runtime confounding is used in several related ways in recent causal-inference and recommender-systems literature. In one formulation, all relevant confounders are captured in historical data, but some cannot be used at prediction time, so treatment assignment is unconfounded given \((V,Z)\) but not given the runtime-available variables \(V\) alone [2006.16916]. In another, a target population lacks confounders \(\mathbf U\) that were observed in the source population, so naively discarding \(\mathbf U\) can lead to severe miscoverage in counterfactual prediction intervals [2604.03772]. In recommender systems, runtime confounding also denotes inference-time bias caused by a confounding feature that directly affects whether an interaction happens, or by a deployed policy that begins to depend on a feature that downstream training still omits [2205.06532]. Taken together, these formulations suggest a deployment-time mismatch between observational training and causal scoring or prediction [2508.10479].

## 1. Core meanings and problem variants

The term has not been restricted to a single formalism. In decision-support settings, runtime confounding refers to the case where historical data contain all relevant factors, but some such factors are unavailable, impermissible, or undesirable in the final prediction model. The paper "Counterfactual Predictions under Runtime Confounding" states this through training ignorability,
\[
Y^a \perp A \mid V, Z,
\]
together with runtime confounding,
\[
Y^a \not\perp A \mid V,
\]
equivalently,
\[
A \not\perp Z \mid V \quad \text{and} \quad Y^a \not\perp Z \mid V
\]
[2006.16916].

In conformal counterfactual prediction, the term denotes a source–target setting in which \(\mathbf V\) is always observed, \(\mathbf U\) is observed only in the source population, and prediction intervals in the target population must depend only on \(\mathbf V\). The target interval \(C_a(\mathbf V)\) is required to satisfy
\[
P(Y(a)\in C_a(\mathbf V)\mid S=0)\ge 1-\alpha,\qquad a\in\mathcal A,
\]
even though \(\mathbf U\) is unavailable in the target population [2604.03772].

In recommender systems, the same phrase has a more operational meaning. A confounding feature \(A\) is an item feature that has a direct effect on \(Y\) independent of true preference; video length is the motivating example in short-video recommendation, because shorter videos are easier to finish even if the user does not like them. If such a feature is used observationally at scoring time, recommendations become biased toward “easy-to-interact” items [2205.06532]. A related deployment-centered account argues that recommender systems can create confounding at runtime when a previously ignored observed feature begins to influence the action policy while downstream estimation still behaves as if that feature were ignorable [2508.10479].

| Setting | Runtime-confounding mechanism | Representative paper |
|---|---|---|
| Counterfactual prediction | Historical confounders unavailable or impermissible at prediction time | [2006.16916] |
| Counterfactual conformal prediction | Source-only confounders missing in the target population | [2604.03772] |
| Causal recommendation | Inference-time scores contaminated by a confounding feature or its induced spurious correlation | [2205.06532] |
| Deployed recommender pipelines | Action policy starts using a feature that later training omits | [2508.10479] |

## 2. Formal causal structure and identification

The counterfactual-prediction formulation is centered on a binary intervention \(A\in\{0,1\}\), runtime-available predictors \(V\), runtime-hidden confounders \(Z\), observed outcome \(Y\), and potential outcomes \(Y^a\). The target is the conditional potential-outcome mean
\[
\nu_a(v) := E[Y^a \mid V=v].
\]
With consistency and positivity, training ignorability implies
\[
\mu_a(v,z) := E[Y^a \mid V=v, Z=z] = E[Y \mid V=v, Z=z, A=a],
\]
and therefore
\[
\nu(v)=E[Y^a\mid V=v] =E\!\left[E[Y\mid V=v,Z=z,A=a]\mid V=v\right] =E[\mu(V,Z)\mid V=v].
\]
The same paper contrasts this target with treatment-conditional regression, which estimates \(E[Y\mid A=a,V=v]\) and incurs pointwise confounding bias
\[
b(v)=\omega(v)-\nu(v)=\int \mu(v,z)\Big(p(z\mid V=v,A=a)-p(z\mid V=v)\Big)\,dz \neq 0
\]
under runtime confounding [2006.16916].

The source–target formulation makes the missing-runtime-confounder structure explicit through the observed unit
\[
\mathbf O_i = (S_iY_i,\ S_iA_i,\ S_i\mathbf U_i,\ \mathbf V_i,\ S_i)\sim P,
\]
with \(Y,A,\mathbf U\) observed only when \(S=1\). Its core assumptions are positivity,
\[
0<P(A=a\mid \mathbf X)<1,
\]
consistency,
\[
Y=\sum_{a\in\mathcal A}\mathbb I(A=a)\,Y(a),
\]
unconfoundedness in the source population,
\[
Y(a)\perp A\mid \mathbf X,\ S=1,
\]
source exchangeability,
\[
Y(a)\perp S\mid \mathbf V,
\]
and source positivity,
\[
0<P(S=1\mid \mathbf V=\mathbf v)<1.
\]
A key point in that formulation is that simply discarding \(\mathbf U\) would require the stronger condition
\[
Y(a)\perp A\mid \mathbf V,\ S=1,
\]
which is not assumed [2604.03772].

The recommender formulation uses a different graph. Item features are split into \(A\), a confounding feature, and \(X\), other content features, with user features \(U\) and interaction label \(Y\). The backdoor path
\[
X \leftarrow Z \rightarrow A \rightarrow Y
\]
implies that learning either \(P(Y\mid U,X)\) or \(P(Y\mid U,X,A)\) from observational data produces biased scoring at inference time. The causal estimand becomes
\[
P(Y \mid U, do(X)) = \sum_{a \in \mathcal{A}} P(Y \mid U, X, A=a)\, P(A=a),
\]
and the runtime correction is
\[
P(y=1 \mid u, do(x)) = \sum_{a \in \mathcal{A}} P(a)\, f(u,x,a).
\]
This shifts the deployed scoring rule from a factual predictor to an interventional predictor [2205.06532].

## 3. Prediction, debiasing, and uncertainty quantification

For counterfactual prediction with restricted runtime covariates, the central methodological response is a two-stage doubly robust procedure. The key pseudo-outcome is
\[
\frac{\mathbb{I}\{A=a\}}{\pi(V,Z)}\left(Y-\mu(V,Z)\right)+\mu(V,Z),
\]
which is then regressed on \(V\) to estimate \(\nu(v)\). The pointwise error bound is product-form:
\[
E \Big[\big(\hat\nu_{\mathrm{DR}}(v)-\nu(v)\big)^2 \Big] \lesssim E\Big[\big(\tilde\nu(v)-\nu(v)\big)^2\Big] + E\Big[\big(\hat\mu(V,Z)-\mu(V,Z)\big)^2 \mid V=v\Big] E\Big[\big(\hat\pi(V,Z)-\pi(V,Z)\big)^2 \mid V=v\Big].
\]
For evaluation, the same paper proposes a doubly robust estimator of the mean squared error of a learned prediction function \(\hat\nu\),
\[
\frac{1}{n}\sum_{i=1}^{n} \left[ \frac{\mathbb{I}\{A_i=a\}}{\hat\pi(V_i,Z_i)} \Big( (Y_i-\hat\nu(V_i))^2-\hat\eta(V_i,Z_i) \Big) +\hat\eta(V_i,Z_i) \right],
\]
and reports in a child-welfare study that treatment-conditional regression had MSE about \(0.290\), the plug-in method about \(0.249\), and the doubly robust method about \(0.248\) [2006.16916].

In runtime-confounded conformal prediction, the objective is not point prediction but valid target-population intervals. The paper "Debiased Machine Learning for Conformal Prediction of Counterfactual Outcomes Under Runtime Confounding" identifies the calibration threshold \(r_{a,\alpha}\) through
\[
P(R_a(Y(a),\mathbf V)\le r_{a,\alpha}\mid S=0)=1-\alpha,
\]
and uses both an identification formula with
\[
E[m_a(r_{a,\alpha},\mathbf V)\mid S=0]=1-\alpha
\]
and a weighted formula with
\[
E\!\left[w_a(\mathbf O)\,\mathbb I(R_a(Y,\mathbf V)\le r_{a,\alpha})\right]=1-\alpha.
\]
Its main estimator is built from the efficient influence curve
\[
\chi_a(r_{a,\alpha},\mathbf O;\eta_a(r_{a,\alpha})),
\]
and solves the empirical estimating equation
\[
P_n\{\chi_a(\hat r_{a,\alpha},\mathbf O;\hat\eta_a(\hat r_{a,\alpha}))\}=0.
\]
Theorem 3 gives the coverage expansion
\[
P(Y(a)\in \hat C_a(\mathbf V)\mid S=0) = 1-\alpha + O_P\!\left(\frac1{\sqrt n}+R_n\right),
\]
where
\[
R_n= \sup_r\|\hat q_a(r,\cdot)-q_a(r,\cdot)\|\cdot\|\hat g_a-g_a\| + \sup_r\|\hat m_a(r,\cdot)-m_a(r,\cdot)\|\cdot\|\hat\kappa-\kappa\|.
\]
The paper states that the naive method miscovers badly at all sample sizes, that miscoverage worsens as runtime confounding becomes more severe, and that the proposed DML method approaches the nominal \(90\%\) coverage as \(n\) grows; the weighted method also achieves near-nominal \(90\%\) coverage, and the proposed DML intervals are often as narrow or narrower than the weighted intervals [2604.03772].

A notable feature of both lines of work is that runtime confounding is treated as a deployment constraint rather than a failure of historical identifiability. The historical data can be rich enough to identify causal structure, while the deployed prediction rule is intentionally restricted.

## 4. Recommender systems and inference-time causal correction

The recommender literature gives runtime confounding a particularly operational interpretation. In "Addressing Confounding Feature Issue for Causal Recommendation," the confounding feature \(A\) directly affects the interaction label \(Y\), so finished interactions do not necessarily indicate preference. The proposed framework, Deconfounding Causal Recommendation (DCR), trains a model to estimate \(P(Y\mid U,X,A)\) but performs recommendation using the interventional quantity \(P(Y\mid U,do(X))\). Direct computation of
\[
\sum_{a \in \mathcal{A}} P(a)\, f(u,x,a)
\]
requires one model evaluation for every possible confounding value \(a\), so if \(|\mathcal A|=K\), inference becomes \(K\) times more expensive. To reduce this cost, DCR introduces a mixture-of-experts architecture with a shared backbone
\[
\mathbf{m} = f_{\Theta}(u,x)
\]
and expert heads
\[
f(u,x,a) = f_{\phi_a}(\mathbf{m} \mid a),
\]
so that
\[
P(y=1 \mid u, do(x)) \approx \sum_{a \in \mathcal{A}} P(a)\, f_{\phi_a}(\mathbf{m} \mid a).
\]
Empirically, with \(K=6\), DCR-MoE achieved the best recommendation accuracy on both datasets. On Kwai it reached Recall@10 \(=0.1089\), MAP@10 \(=0.0353\), and NDCG@10 \(=0.0634\); on Wechat it reached Recall@3 \(=0.1355\), MAP@3 \(=0.0976\), and NDCG@3 \(=0.1271\). Reported inference times were \(13.6\)s and \(2.9\)s for DCR-NFM, \(5.4\)s and \(1.1\)s for DCR-MoE, and \(2.5\)s and \(0.6\)s for the approximation-based DCR-NFM-A on Kwai and Wechat respectively [2205.06532].

A second recommender formulation emphasizes system evolution rather than static item features. The paper "Confounding is a Pervasive Problem in Real World Recommender Systems" uses variables \(c\) for click outcome, \(a\) for recommended action, \(x_1\) for the feature currently used for personalization, and \(x_2\) for an additional feature that may later be introduced. The causal target is
\[
P(c \mid {\rm do}(a), x_1) = \sum_{x_2} P(c \mid a, x_1, x_2) P(x_2 \mid x_1).
\]
If \(x_2\) does not affect the action, this simplifies to
\[
P(c \mid {\rm do}(a), x_1) = P(c \mid a, x_1).
\]
But once the policy starts using \(x_2\), and later training still omits it, \(x_2\) becomes a confounder. The paper describes this as a temporal sequence: a randomized policy on Day 0, a policy using \(x_1\) on Day 1, a policy using both \(x_1\) and \(x_2\) on Day 2, and then a reversion on Day 3 to training with only \(x_1\) on logs generated by a policy that depended on \(x_2\). It identifies feature engineering, A/B testing on shared logs, and modularization as mechanisms that can create runtime confounding in deployed systems [2508.10479].

A common misconception is that recommender systems are safe from confounding because all inputs are “observed.” The recommender papers explicitly reject that view: an observed feature can become a confounder when it influences both action selection and outcome, and deployment changes can alter the causal graph without changing the training code [2508.10479].

## 5. Relation to broader confounding methodologies

Runtime confounding sits within a broader literature on causal inference under confounding, but it is not reducible to any one classical problem. The instrumental-variables literature addresses a different obstacle: confounding by an unmeasured \(U\) when ordinary regression fails. In the simple structural equation
\[
Y = A\beta + U,
\]
ordinary least squares gives
\[
\hat{\beta}_{OLS} = \beta + (A^TA)^{-1}A^T U,
\]
which is consistent only if \(E[A^T U]=0\). An instrumental variable \(Z\) satisfying relevance, independence, and exclusion restriction yields
\[
\hat{\beta}_{IV} = (Z^T A)^{-1} Z^T Y,
\]
and the appendix generalizes this to two-stage least squares [2506.18652]. This does not solve runtime confounding directly, but it addresses the adjacent case where confounders are not observed at all.

Safe decision-making under hidden confounding leads to yet another response. "Confounding-Robust Policy Improvement" assumes that policy value and regret may not be point-identifiable under unobserved confounding and therefore optimizes worst-case regret relative to a baseline policy \(\pi_0\). The method uses a marginal sensitivity model with odds-ratio bound
\[
\Gamma^{-1}\leq \frac{(1-\tilde e_T(X))\,e_T(X,Y)} {\tilde e_T(X)\,(1-e_T(X,Y))} \leq \Gamma,
\]
and learns a policy by minimizing worst-case empirical regret over an uncertainty set of inverse propensity weights. The paper emphasizes safety relative to baseline rather than point identification of a fully personalized policy [1805.08593].

Observed-confounding conformal prediction provides a finite-sample back-door analogue. "Conformal e-prediction in the presence of confounding" studies the graph
\[
Z \to X,\qquad X \to Y,\qquad Z \to Y,
\]
and targets the interventional law
\[
p_y := \sum_{z\in\mathbf Z} P(Z=z)\,P(Y=y\mid X=x,Z=z).
\]
It constructs a smoothed estimator \(F_y\) from empirical counts and proves
\[
\mathbb E\frac{p_y}{F_y}\le 1,
\]
which yields e-values and prediction regions for \(Y\) under \(do(X=x)\) [2603.11134]. This is not a runtime-confounding paper in the narrow sense, but it clarifies how prediction targets change once one moves from \(P(Y\mid X=x)\) to \(P(Y\mid do(X=x))\).

Causal discovery under confounding addresses a distinct but related question. LiNGAM-MMI replaces the standard LiNGAM requirement that one order achieve independent residuals with the objective
\[
\arg\min_{\text{order}}\, K(e_1,\ldots,e_p),
\]
where
\[
K(e_1,\ldots,e_p) :=E\!\left[\log\frac{P(e_1,\ldots,e_p)}{P(e_1)\cdots P(e_p)}\right].
\]
The method interprets larger residual dependence as stronger confounding and searches for the globally optimal order by a shortest-path formulation [2401.16661]. This is adjacent to runtime confounding insofar as it treats confounding-aware causal structure as a prerequisite for later deployment.

## 6. Limitations, misconceptions, and practical significance

Several misconceptions recur across the literature. Runtime confounding is not the same as standard confounding in a single population, because the defining issue is often that causal adjustment is possible in training data but not in the deployed predictor [2006.16916]. It is also not the same as target shift or a generic missing-covariate problem; the source–target conformal paper states that the key issue is that some confounders are available in training but not at runtime in the target site, and it attributes the problem to two simultaneous shifts: covariate shift across treatment levels within the source population and covariate shift between source and target populations in \(\mathbf V\) [2604.03772]. Nor is runtime confounding restricted to unmeasured causes: in recommender systems, ignored observed features can become confounders when policy changes make them influence actions [2508.10479].

The limitations are equally consistent. Runtime-confounding corrections often impose computational or modeling costs. In DCR, exact backdoor adjustment requires summing over all confounder values, which creates the runtime bottleneck that motivates the mixture-of-experts architecture; the approximation-based alternative is fastest but sacrifices accuracy [2205.06532]. In conformal DML, a full conformal version without data splitting is possible but requires stronger Donsker-type conditions and is computationally heavier [2604.03772]. In IV-based inference, independence and exclusion restriction are not directly testable when confounders are unmeasured, so identification remains fundamentally a matter of theory and substantive knowledge [2506.18652]. In sensitivity-based policy learning, larger \(\Gamma\) gives stronger protection against hidden confounding but can be conservative if the real confounding is smaller [1805.08593].

The practical significance is that runtime confounding converts an apparently ordinary prediction problem into a causal transport-and-deployment problem. Naive treatment-conditional regression can target the wrong quantity even when fit perfectly [2006.16916]. Naively dropping source-only confounders can break interval validity [2604.03772]. Naively using observational recommender scores can over-recommend short videos or otherwise exploit “easy-to-interact” confounding values [2205.06532]. Naively pooling logs across feature-mismatched recommender variants can entrench bias in A/B testing and modularized systems [2508.10479]. These results suggest that runtime confounding is best understood not as a narrow technical anomaly, but as a recurring mismatch between the variables that support causal identification during learning and the variables that remain available, admissible, or consistently used when decisions are made.

Source: https://www.emergentmind.com/topics/runtime-confounding