---
title: Evidence-Supported Bounds in Causal Inference
url: https://www.emergentmind.com/topics/evidence-supported-bounds
type: topic
---

# Evidence-Supported Bounds in Causal Inference

Searching arXiv for recent and foundational papers on “evidence-supported bounds” and closely related formulations across causal inference, evaluation, and uncertainty quantification.
Evidence-supported bounds are interval-valued conclusions that delimit the set of parameter values, performance levels, probabilities, or effects that remain compatible with the available evidence and stated assumptions. Across the literature, they arise when point identification is unavailable, undesirable, or epistemically misleading. In causal inference, they appear as partial-identification statements for quantities such as the probability of causation or the population average treatment effect [1411.2636; 1801.07328]. In benchmark evaluation, they quantify uncertainty induced by incomplete retained artifacts [2605.10448]. In review and verification systems, closely related label sets encode how far available evidence carries a claim [2604.04074]. In each setting, the common logic is to replace unsupported point assertions with ranges or graded judgments whose width reflects unresolved uncertainty rather than sampling variation alone.

## 1. Evidence-supported bounds as partial identification

Evidence-supported bounds are most explicit in work that treats inference as a partial-identification problem. In the generalization of randomized trials without probability sampling, bounding is defined as estimating a range of plausible values for the population treatment effect rather than a single point estimate, because the population average treatment effect cannot generally be uniquely identified from the data alone when sample selection is nonignorable [1801.07328]. The same paper states that the framework nonparametrically estimates population parameters using a range of plausible values that are consistent with the observed characteristics of the data, and that this range is driven by the size of the nonsampled share of the population and the permissible outcome range [1801.07328].

A similar logic governs the probability of causation. The quantity
\[
\pc_A \;=\; \Pr_A\!\big(Y_A(0)=0 \mid X_A=1,\, Y_A(1)=1\big)
\]
is an individual-level counterfactual probability, but the joint dependence between \(Y(0)\) and \(Y(1)\) is not observed, so only bounds are generally available [1411.2636]. The standard empirical-dependence bounds are
\[
1 - \frac{1}{RR} \;\leq\; \pc_A \;\leq\; \frac{\Pr(Y=0 \mid X=0)}{\Pr(Y=1 \mid X=1)},
\]
with
\[
RR = \frac{\Pr(Y=1 \mid X=1)}{\Pr(Y=1 \mid X=0)}.
\]
The paper characterizes these as “group-to-individual” bounds because they use only the empirical dependence between exposure and outcome and make no assumptions about the dependence between the two potential outcomes \(Y(0)\) and \(Y(1)\) [1411.2636].

This partial-identification perspective also appears in discrete-choice panel data with individual-specific effects. There, the sharp identified set for average effects is practically infeasible at realistic sample sizes whenever the number of support points of the observed covariates is large, and the proposed alternative is to estimate outer bounds on the identified set [2309.09299]. These bounds satisfy
\[
\mathbb E[L(Z_i,Y_i,\beta)] \le \overline m \le \mathbb E[U(Z_i,Y_i,\beta)],
\]
and are described as easy to construct, converging at the parametric rate, and computationally simple to obtain even in moderately large samples [2309.09299].

A plausible implication is that “evidence-supported bounds” functions less as a domain-specific term than as a general inferential stance: conclusions should contract only to the extent justified by identifiable structure, observed margins, and explicitly adopted assumptions.

## 2. Causal inference: probability of causation, mediation, and principal effects

In causal inference, evidence-supported bounds are used to address counterfactual questions that cannot usually be identified exactly from population data. For the probability of causation, the 2014 mediation paper shows how the standard bounds can be adapted or improved if further information becomes available, including pre-treatment covariates, observational data under confounding, and a mediator observed in the experiment [1411.2636].

With only experimental evidence on binary exposure \(X\) and outcome \(Y\), the baseline bounds are the standard risk-ratio bounds above. The paper gives an aspirin-like example with
\[
\Pr(Y=1\mid X=1)=0.30,\qquad \Pr(Y=1\mid X=0)=0.12,
\]
so that \(RR=2.5\), yielding
\[
0.60 \le \pc_A \le 1.
\]
It notes that the lower bound exceeds \(1/2\) whenever \(RR>2\), which is often the civil-law threshold for “more likely than not” [1411.2636].

When a pre-treatment covariate \(S\) is available, the same analysis can be performed conditionally on \(S=s_A\), using
\[
RR(s)=\frac{\Pr(Y=1\mid X=1,S=s)}{\Pr(Y=1\mid X=0,S=s)}.
\]
If \(S\) is only observed in the data and not for the individual, refined bounds are available through the quantities \(\Delta\) and \(\Gamma\), and the paper states that these refined bounds are never wider than the unadjusted bounds [1411.2636].

The mediation extension isolates a binary mediator \(M\) on the path \(X \to M \to Y\), under no direct effect and no confounding among the mechanisms. In that regime, the lower bound is unchanged from the no-mediator case, but the upper bound can improve [1411.2636]. The paper’s explicit conclusion is that mediator information is useful for tightening the upper bound on causation probability, while the lower bound is unaffected. In its example,
\[
\Pr(M=1\mid X=1)=0.25,\qquad \Pr(M=1\mid X=0)=0.025,
\]
\[
\Pr(Y=1\mid M=1)=0.9,\qquad \Pr(Y=1\mid M=0)=0.1,
\]
yield
\[
0.60 \le \pc_A \le 0.76,
\]
which is sharper than the unmediated analysis, where the upper bound was vacuous [1411.2636].

The 2017 follow-up extends this line to partial mediation [1706.04857]. It defines \(Y^*(x,m)\) and uses assumptions
\[
Y^*(x,m) \perp M \mid X,\qquad Y^*(x,m) \perp X,\qquad M(x) \perp X.
\]
The paper derives mediator-based upper bounds by decomposing the joint event \(Y(0)=0, Y(1)=1\) over the four mediator potential-outcome combinations \(m_0,m_1\in\{0,1\}\), using the inequality
\[
\max\{P(A)+P(B)-1,\,0\} \le P(A\cap B)\le \min\{P(A),P(B)\}.
\]
Its substantive comparison is precise: in complete mediation, incorporating the mediator never worsens the upper bound relative to ignoring mediation, whereas in partial mediation the mediator-based upper bound may be smaller or larger, so both should be computed and compared [1706.04857].

Bounding also appears in principal stratification. The Early College High Schools analysis argues that subgroup effects defined by latent post-treatment behavior are often not point identified, and that bounds can provide the set of all values consistent with the observed data and assumptions [1701.03139]. Under monotonicity, irrelevant alternatives, and an exclusion restriction for always-takers, the paper derives sharp bounds for complier-specific effects and shows that baseline covariates can substantially tighten them [1701.03139]. The central empirical result is that the program’s impact on ninth-grade “on-track” status is substantial for students who would otherwise attend a low-quality high school, but not for those who would otherwise attend a high-quality high school [1701.03139].

## 3. Tightening strategies: assumptions, covariates, structure, and auxiliary evidence

A recurring feature of evidence-supported bounds is that width is not fixed: it is reduced by additional structure that remains weaker than full identifying assumptions. The generalization study presents three such strategies—monotonicity, redefining the population of inference, and propensity score stratification [1801.07328]. In worst-case form, with outcomes bounded in \([Y_L,Y_U]\), the population average treatment effect satisfies
\[
\text{PATE} \in [A_L, A_U]
\]
with
\[
A_L = \text{SATE}\Pr(Z=1) + (Y_L-Y_U)\Pr(Z=0), \qquad
A_U = \text{SATE}\Pr(Z=1) + (Y_U-Y_L)\Pr(Z=0).
\]
The width is
\[
2(Y_U-Y_L)\Pr(Z=0),
\]
which the paper emphasizes is often uninformatively wide [1801.07328].

The first tightening device is monotone sample selection (MSS), which assumes that schools participating in the trial do so because they expect a larger treatment benefit than nonparticipants. Under MSS, the upper bound improves to \(A_U=\text{SATE}\), while the lower bound is unchanged [1801.07328]. The second device is redefining the population to a more comparable subgroup, which changes the estimand rather than merely the estimator. The third is propensity score stratification, used not for point identification under sampling ignorability, but to compute and average stratum-specific bounds [1801.07328]. Simulation findings are explicit: MSS reduced width by about 40% on average; more restrictive population redefinitions yielded about 20% width reduction on average; and combining MSS with population redefinition could reduce width by nearly 50% relative to the original worst-case bounds [1801.07328].

The principal-stratification paper gives an analogous account of covariate-based tightening. It distinguishes “principal variables,” which predict stratum membership, from “prognostic variables,” which predict outcomes [1701.03139]. The simulations show that principal variables are most beneficial because they separate the unidentified complier types, while prognostic variables also help by pushing slice-specific means toward 0 or 1, where truncation becomes binding [1701.03139]. Compliance variables are generally less useful in that setup because the crucial uncertainty is between two complier types rather than between compliers and always-takers [1701.03139].

In short-panel discrete-choice models, the tightening strategy is algorithmic rather than substantive. The outer bounds are constructed from functions \(L\) and \(U\) satisfying
\[
\sum_{y\in\mathcal Y} L(z,y,\beta)\, f(y\mid z,a;\beta) \;\le\; m(z,a,\beta) \;\le\; \sum_{y\in\mathcal Y} U(z,y,\beta)\, f(y\mid z,a;\beta)
\]
for every \(z\) and \(a\) [2309.09299]. The paper stresses that the optimization is over the domain of \(A\) only, not over the full unknown distribution \(\pi(a|z)\), which is the main reason the procedure is computationally simple even with continuous covariates [2309.09299].

This suggests a common pattern: bounds are tightened either by adding credible constraints on the data-generating mechanism or by reparameterizing the inference problem so that only directly supportable inequalities need to be estimated.

## 4. Evaluation, auditing, and claim verification under incomplete evidence

In interactive-agent evaluation, evidence-supported bounds are operationalized directly as benchmark score intervals. The 2026 benchmark paper introduces an outcome-evidence reporting layer on top of existing benchmarks without changing tasks, agents, environments, or native evaluators [2605.10448]. For each completed run, a locked checklist is applied to retained traces, logs, state snapshots, receipts, and outputs, and one of three evidence labels is assigned: **Evidence Pass**, **Evidence Fail**, or **Unknown** [2605.10448]. A record is Unknown only when the retained artifacts are insufficient to decide the benchmark’s own success claim [2605.10448].

Let \(P\), \(F\), and \(U\) denote the number of Evidence Pass, Evidence Fail, and Unknown records, with \(N=P+F+U\). The formal definitions are
\[
\mathrm{CountedScore}(a,d)=\frac{P}{P+F},\quad
\mathrm{Perf}(a,d)\in\left[\frac{P}{N},\frac{P+U}{N}\right],\quad
\mathrm{Width}(a,d)=\frac{U}{N}.
\]
The lower bound treats every Unknown as a failure; the upper bound treats every Unknown as a success; and the width is the fraction of completed records whose outcome is not identified by the retained evidence [2605.10448]. The paper explicitly classifies these as partial-identification bounds, not as statistical confidence intervals [2605.10448].

A central methodological claim is that counting Unknown as success overstates performance, counting Unknown as failure punishes missing instrumentation rather than agent behavior, and dropping Unknown creates post-selection bias because different agents leave different amounts of evidence [2605.10448]. The alternative is to keep Unknown in the denominator and report an interval over all completed records [2605.10448]. The benchmark reports across ANDROIDWORLD, AGENTDOJO, APPWORLD, \(\tau^3\)-bench retail, and MINIWOB reveal distinct failure modes, including evidence gaps, native evaluator conflicts, reward/action mismatch, and weaker-than-task-wording evaluators [2605.10448].

A related but non-interval formulation appears in evidence-grounded review. FactReview treats peer review as claim verification under evidence limits and assigns each major claim one of five labels: **Supported**, **Supported by the paper**, **Partially supported**, **In conflict**, or **Inconclusive** [2604.04074]. It does not provide a symbolic score-bounding rule, but the paper repeatedly frames the method as shrinking claims to fit evidence. The CompGCN case study is exemplary: reproduced results closely match reported values for FB15k-237 link prediction and MUTAG node classification, but on MUTAG graph classification the reproduced result is \(88.4\%\) while the strongest baseline reported in the paper remains \(92.6\%\), so the broad “outperforms across tasks” claim is only partially supported [2604.04074].

A plausible implication is that interval bounds and graded support labels instantiate the same epistemic policy. Where a scalar target can be partially identified, one reports a range. Where claim scope is heterogeneous or multidimensional, one reports support categories tied to the evidence sources that do or do not sustain each subclaim.

## 5. Beyond causal and evaluation settings: evidence, support, and uncertainty quantification

The phrase “evidence-supported bounds” also appears in settings where the object being bounded is not a causal effect or benchmark score.

In Bayesian model comparison for singular statistical models, the object is the marginal likelihood or evidence. The singular-model paper writes
\[
m(X^{(n)})=\int_{\Omega} e^{\ell_n(\xi)}\varphi(\xi)\,d\xi
\]
and studies
\[
Z(n)=\int_\Omega e^{-nK_n(\xi)}\varphi(\xi)\,d\xi.
\]
In normal form, with
\[
K(\xi)=\xi^{2k}, \qquad \varphi(\xi)=b(\xi)\,\xi^h,
\]
the real log canonical threshold is
\[
\lambda=\min_j \frac{h_j+1}{2k_j}, \qquad
m=\#\left\{j:\frac{h_j+1}{2k_j}=\lambda\right\}.
\]
The paper proves a non-asymptotic two-sided bound
\[
C_1 \frac{(\log n)^{m-1}}{n^\lambda} \;<\; Z_K(n) \;<\; C_2 \frac{(\log n)^{m-1}}{n^\lambda},
\]
thereby recovering the correct evidence scaling in \(n\) and \(\log n\) without invoking the usual regular-model parameter counting [2008.04537]. A new Gibbs-variational bound further shows that mean-field variational inference recovers the RLCT rate, in the sense that
\[
\sup_{\rho\in\mathcal F_{\mathrm{MF}}}\Psi_n(\rho)\ge -\lambda\log n - C
\]
for the mean-field family \(\mathcal F_{\mathrm{MF}}\) [2008.04537].

In evidential statistics, support intervals invert Bayes factors. With approximately normal estimator \(\hat\theta \sim \Nor(\theta,\sigma^2)\), the support interval is
\[
\mathrm{SI}_k = \left\{\theta_0 : BF_{01}(x;\theta_0)\ge k\right\},
\]
meaning that the observed data are at least \(k\) times more likely under the included parameter values than under a specified alternative [2206.12290]. The paper also derives a coverage bound from the universal Bayes-factor bound:
\[
\Pr(\theta_* \in \mathrm{SI}_k \mid \theta=\theta_*) \ge 1-k,\qquad k<1.
\]
This connects evidence-based support directly to a lower bound on repeated-sampling coverage [2206.12290].

In Bayesian networks, sensitivity analysis can be made evidence-invariant. Under proportional co-variation, any sensitivity function has the form
\[
f(x)=\frac{C_3x+C_4}{C_1x+C_2}
\]
or equivalently
\[
f(x)=\frac{r}{x-s}+t.
\]
The evidence-invariant paper gives envelope bounds on such functions and the derivative bound
\[
|f'(x_0)| \le \frac{p_0(1-p_0)}{x_0(1-x_0)}
\]
for the sensitivity value at the original parameter value [1207.4170]. The follow-up shows that once the constant \(s=C_2/C_1\) is known from the probability of the evidence, tighter evidence-dependent bounds on both slope and vertex location become available [1207.1357]. These bounds support screening parameters whose local sensitivity is small but whose vertex lies nearby, so that modest finite perturbations could still induce large changes [1207.1357].

In probability-of-evidence computation for belief networks, the object being bounded is itself the evidence probability \(P(e)\). The Markov-LB method uses importance sampling and Markov’s inequality to produce a high-confidence lower bound of the form
\[
\frac{f(x)}{a\,Q(x)} \le P(e)
\quad\text{with probability at least }1-\frac1a,
\]
and with \(k\) independent samples, the minimum bound is valid with probability at least \(1-\frac{1}{a^k}\) [1206.5242]. The paper emphasizes that these are high-confidence lower bounds, not relative- or absolute-error approximations [1206.5242].

In epistemic uncertainty propagation using Dempster–Shafer structures on closed intervals, evidence-supported bounds are the induced output focal-element intervals obtained by propagating each input focal element through a polynomial-chaos surrogate and a Bernstein range enclosure [1107.1547]. For each focal-element pair, the output interval is
\[
Y_{ij}=[\underline{y}_{ij},\overline{y}_{ij}], \qquad
\underline{y}_{ij}=\min \beta_{\mathbf I}, \quad
\overline{y}_{ij}=\max \beta_{\mathbf I},
\]
where the Bernstein coefficients bound the polynomial surrogate by the range enclosure theorem [1107.1547].

## 6. Interpretation, misconceptions, and methodological significance

A persistent misconception is to conflate evidence-supported bounds with statistical confidence intervals. The benchmark-evaluation paper is explicit that its intervals represent the set of all all-record performance values still compatible with the evidence, not sampling uncertainty [2605.10448]. The causal-inference papers make the same distinction implicitly: the width of the bounds comes from unobserved counterfactual dependence, latent stratum membership, or nonsampled population segments, not from finite-sample variance alone [1411.2636; 1701.03139; 1801.07328]. Confidence intervals may be layered on top of such bounds, as in the panel-data paper’s asymptotically valid confidence intervals for the identified set or outer bounds [2309.09299], but the underlying inferential object remains a set identified by assumptions and observed evidence.

A second misconception is that wider bounds are method failures. Several papers argue the opposite. Wide bounds can reveal that the available artifacts do not decide benchmark outcomes [2605.10448], that the trial sample does not support extrapolation to the full target population [1801.07328], or that mediator information only partially resolves cross-world uncertainty [1706.04857]. In this sense, width is itself a measurement of unresolved epistemic uncertainty.

A third misconception is that more structure always helps. The mediation literature shows that complete mediation can only improve or preserve the upper bound relative to ignoring mediation, but partial mediation may yield an upper bound that is smaller or larger than the non-mediation bound [1706.04857]. The generalization paper similarly warns that tighter bounds are valuable only if the added assumption is substantively credible; when MSS was weakly plausible or false, coverage fell sharply, sometimes below 25% [1801.07328]. This suggests that narrowing a bound by adding assumptions is not automatically an evidential gain unless those assumptions are themselves well supported.

The broader methodological significance is that evidence-supported bounds provide a disciplined alternative to unsupported point summaries. They preserve the distinction between what is observed, what is assumed, and what remains unresolved. Across causal inference, benchmark evaluation, sensitivity analysis, Bayesian evidence, and uncertainty propagation, the same principle recurs: when the evidence does not justify a single value, the scientifically appropriate output is a constrained set, an interval, or a graded support judgment whose scope tracks the evidence exactly [1411.2636; 2605.10448; 2008.04537].

Source: https://www.emergentmind.com/topics/evidence-supported-bounds