Papers
Topics
Authors
Recent
Search
2000 character limit reached

Surrogate Index in Causal Inference

Updated 10 July 2026
  • Surrogate index is a method that condenses multiple short-term surrogate measures into a composite estimator for long-term treatment effects.
  • It employs techniques like penalized regression, tree-based models, and ensemble methods to manage high-dimensional surrogate data effectively.
  • Extensions include proximal and weighted approaches to address unobserved confounding, improve decision-making in experiments, and enable trial-level meta-analysis.

A surrogate index is a construction that compresses surrogate information into an estimand, predictor, or composite marker used in place of an unavailable target. In causal inference, the canonical surrogate index is the conditional mean of a long-term outcome given short-term surrogates and pre-treatment covariates, estimated in one sample and transported to another (Athey et al., 2016). Recent work extends this formulation to settings with unobserved confounding through the proximal surrogate index, defined as an outcome bridge involving proxy variables (Hung et al., 25 Jan 2026), and to sensitivity analysis through Weighted Surrogate Indices that relax full mediation via a conditional copula (Fan et al., 28 Feb 2026). In high-dimensional clinical-trial meta-analysis, a surrogate index may instead be a rank-based composite signature formed from promising trial-level markers (Hughes et al., 5 May 2026). In surrogate modeling, the phrase is also used for low-dimensional projections or admissible multi-index sets that organize emulator structure and adaptive refinement (Li et al., 2023, Kent et al., 4 Jul 2025).

1. Classical surrogate index for long-term treatment effects

The foundational causal-inference formulation considers two samples. The experimental sample contains treatment W{0,1}W\in\{0,1\} and short-term surrogates SRdS\in\mathbb R^d but no long-term outcome YY, while the observational sample contains SS and YY but no WW. The surrogate index is defined as

μ(s,x,p)  :=  E[YS=s,X=x,P=p],\mu(s,x,p)\;:=\;E[Y\mid S=s,X=x,P=p],

with P{E,O}P\in\{E,O\} denoting the sample indicator. Under comparability, μ(s,x,E)=μ(s,x,O)\mu(s,x,E)=\mu(s,x,O), so μ(s,x,O)\mu(s,x,O) estimated in the observational sample can be used to impute long-term outcomes in the experiment (Athey et al., 2016).

Identification relies on three core conditions. First, unconfoundedness in the experiment requires SRdS\in\mathbb R^d0 with overlap. Second, Prentice-type surrogacy requires SRdS\in\mathbb R^d1 together with overlap in SRdS\in\mathbb R^d2. Third, comparability requires SRdS\in\mathbb R^d3 and support inclusion of SRdS\in\mathbb R^d4 in SRdS\in\mathbb R^d5. These conditions imply that treatment affects the long-term outcome only through the observed surrogates once SRdS\in\mathbb R^d6 are held fixed, and that the conditional law of SRdS\in\mathbb R^d7 given SRdS\in\mathbb R^d8 is transportable across the two samples (Athey et al., 2016).

Under these assumptions, the average treatment effect in the experiment,

SRdS\in\mathbb R^d9

can be written as

YY0

An equivalent inverse-probability-weighted representation is

YY1

where YY2. Operationally, the estimator is therefore a difference in means on an imputed long-term outcome (Athey et al., 2016).

A central motivation is that multiple surrogates may collectively satisfy the statistical surrogacy criterion even when no individual proxy does so alone. This is particularly relevant in modern settings with large numbers of intermediate outcomes on or near the causal pathway between treatment and a long-term endpoint (Athey et al., 2016).

2. Estimation with multiple and high-dimensional surrogates

When YY3 is high-dimensional, the surrogate index is estimated by regressing YY4 on YY5 in the observational sample. The methods explicitly listed for this step include penalized linear estimators such as Lasso, elastic net, and ridge; tree-based estimators such as random forests, boosted trees, and Bayesian additive regression trees; and “Super-Learner” ensembles. A generic procedure is: fit YY6 in the observational sample, compute YY7 in the experimental sample, and estimate the treatment effect by difference in means, weighted regression, or IPW with YY8 (Athey et al., 2016).

Cross-fitting may be used to remove over-fitting bias. The influence-function approach yields a “double-robust” estimator that remains YY9-consistent if either SS0 or SS1 is well-estimated. This places surrogate-index estimation within the general semiparametric program of nuisance-robust estimation with flexible first-stage learners (Athey et al., 2016).

The California GAIN application illustrates the approach empirically. There, the primary outcome is average quarterly employment rate or earnings over 36 quarters; the surrogates are employment, earnings, and AFDC receipt in the first SS2 quarters; and the observational sample consists of other sites while the experimental sample is the Riverside site. A surrogate index based on SS3 quarters reproduces the 36-quarter ATE on employment, reported as approximately SS4 percentage points, and does so about three years earlier. Standard errors shrink by roughly SS5, with the example given as SS6 percentage points to SS7 percentage points. By contrast, naive short-window estimators remain biased until about SS8, and tests for surrogacy and comparability indicate that SS9 quarters suffice to remove statistically significant direct YY0–YY1 associations (Athey et al., 2016).

These results indicate that the surrogate index is not merely a dimensionality-reduction device. It is an identification-driven summary: the projection of a possibly high-dimensional surrogate vector onto the conditional expectation most relevant for the long-term outcome.

3. Experimentation and decision-making with auto-surrogate indices

In large-scale online experimentation, the surrogate index has been used as a decision tool rather than only as an ATE estimator. A Netflix study evaluates linear “auto-surrogate” models that use the shorter-term observations of the long-term outcome itself. For a user YY2, the target long-term quantity is

YY3

and the linear auto-surrogate predictor based on the first YY4 days is

YY5

The coefficients are estimated by OLS (Zhang et al., 2023).

The empirical setting contains 200 A/B tests and 1,098 non-control treatment arms in Netflix’s personalization space, with treatment arms and controls randomly assigned by “single-shot” allocation and typically tens or hundreds of thousands of users per arm. Two training regimes are considered for YY6: a pre-test regime that trains OLS on each user’s 63-day history immediately prior to allocation, and a similar-test regime that trains on a different but similar A/B test in the same product area (Zhang et al., 2023).

The decision rule maps each arm to YY7 or YY8, depending on whether the estimated treatment effect is positive and statistically significant at level YY9, negative and significant, or otherwise nonsignificant. The principal evaluation criterion is consistency,

WW0

together with precision and recall for positive launches (Zhang et al., 2023).

Using 14 days of data to infer day-63 effects, the reported overall decision consistency is approximately WW1. The surrogate produces nonsignificant reads more often, with WW2 versus WW3. Restricting attention to the set of tests that would be launched according to the 63-day directly measured treatment effects, the surrogate-based decision achieves precision of approximately WW4 and recall of approximately WW5. The report also states that there are no false-positive launches: the surrogate never indicates “launch” when the direct estimate is significantly negative. Under additivity and no marginal cost, shifting from 63-day to 14-day decisions could increase experiment capacity by up to WW6, while WW7 recall implies that about WW8 more 14-day experiments would be needed to match the gains of the 63-day cycle (Zhang et al., 2023).

The assumptions are explicit: the short-term metrics must fully mediate the treatment effect on the long-term outcome, there must be no unmeasured confounders linking treatment, WW9, and μ(s,x,p)  :=  E[YS=s,X=x,P=p],\mu(s,x,p)\;:=\;E[Y\mid S=s,X=x,P=p],0, and the linear model with i.i.d. errors must be correctly specified. The same study emphasizes model misspecification, training-data drift, and fat-tailed treatment-effect distributions as practical risks, and recommends periodic validation on holdout long-term tests (Zhang et al., 2023).

4. Failures of surrogacy, proximal identification, and sensitivity analysis

The classical surrogate-index approach is vulnerable when Prentice surrogacy fails. If μ(s,x,p)  :=  E[YS=s,X=x,P=p],\mu(s,x,p)\;:=\;E[Y\mid S=s,X=x,P=p],1, then the surrogate-index estimator converges not to the true ATE but to

μ(s,x,p)  :=  E[YS=s,X=x,P=p],\mu(s,x,p)\;:=\;E[Y\mid S=s,X=x,P=p],2

where μ(s,x,p)  :=  E[YS=s,X=x,P=p],\mu(s,x,p)\;:=\;E[Y\mid S=s,X=x,P=p],3. The bias depends on the residual dependence of μ(s,x,p)  :=  E[YS=s,X=x,P=p],\mu(s,x,p)\;:=\;E[Y\mid S=s,X=x,P=p],4 on μ(s,x,p)  :=  E[YS=s,X=x,P=p],\mu(s,x,p)\;:=\;E[Y\mid S=s,X=x,P=p],5 given μ(s,x,p)  :=  E[YS=s,X=x,P=p],\mu(s,x,p)\;:=\;E[Y\mid S=s,X=x,P=p],6 and on μ(s,x,p)  :=  E[YS=s,X=x,P=p],\mu(s,x,p)\;:=\;E[Y\mid S=s,X=x,P=p],7. The same analysis notes that the bias is small if μ(s,x,p)  :=  E[YS=s,X=x,P=p],\mu(s,x,p)\;:=\;E[Y\mid S=s,X=x,P=p],8 explains little of the variation in μ(s,x,p)  :=  E[YS=s,X=x,P=p],\mu(s,x,p)\;:=\;E[Y\mid S=s,X=x,P=p],9 or if P{E,O}P\in\{E,O\}0 explains a large variation in P{E,O}P\in\{E,O\}1; for binary P{E,O}P\in\{E,O\}2, a sharp bound is obtained by replacing the residual term with its maximal possible value P{E,O}P\in\{E,O\}3 or P{E,O}P\in\{E,O\}4 (Athey et al., 2016).

The proximal surrogate index addresses a different failure mode: unobserved confounding. The data combine an experimental sample P{E,O}P\in\{E,O\}5, where P{E,O}P\in\{E,O\}6 are observed but P{E,O}P\in\{E,O\}7 is missing, with an observational sample P{E,O}P\in\{E,O\}8, where P{E,O}P\in\{E,O\}9 are observed but μ(s,x,E)=μ(s,x,O)\mu(s,x,E)=\mu(s,x,O)0 is missing. An unobserved confounder μ(s,x,E)=μ(s,x,O)\mu(s,x,E)=\mu(s,x,O)1 affects μ(s,x,E)=μ(s,x,O)\mu(s,x,E)=\mu(s,x,O)2. The proxy variables satisfy

μ(s,x,E)=μ(s,x,O)\mu(s,x,E)=\mu(s,x,O)3

and the target is

μ(s,x,E)=μ(s,x,O)\mu(s,x,E)=\mu(s,x,O)4

Under no direct effect of μ(s,x,E)=μ(s,x,O)\mu(s,x,E)=\mu(s,x,O)5 on μ(s,x,E)=μ(s,x,O)\mu(s,x,E)=\mu(s,x,O)6 beyond μ(s,x,E)=μ(s,x,O)\mu(s,x,E)=\mu(s,x,O)7, transportability, overlap, and completeness conditions, one seeks an outcome bridge μ(s,x,E)=μ(s,x,O)\mu(s,x,E)=\mu(s,x,O)8 satisfying

μ(s,x,E)=μ(s,x,O)\mu(s,x,E)=\mu(s,x,O)9

This function is called the proximal surrogate index. It identifies μ(s,x,O)\mu(s,x,O)0 through

μ(s,x,O)\mu(s,x,O)1

and hence identifies the experimental-population ATE (Hung et al., 25 Jan 2026).

The same framework derives a surrogate-bridge representation through functions μ(s,x,O)\mu(s,x,O)2 and combines both views in a single doubly robust representation μ(s,x,O)\mu(s,x,O)3. Estimation is multiply robust and cross-fitted: it estimates the experimental propensity μ(s,x,O)\mu(s,x,O)4, the outcome bridge μ(s,x,O)\mu(s,x,O)5, the regression of μ(s,x,O)\mu(s,x,O)6 on μ(s,x,O)\mu(s,x,O)7, and the surrogate bridges μ(s,x,O)\mu(s,x,O)8, then averages held-out influence scores. Under stated μ(s,x,O)\mu(s,x,O)9-error and product-rate conditions,

SRdS\in\mathbb R^d00

with the empirical variance of the foldwise scores providing a consistent variance estimator (Hung et al., 25 Jan 2026).

In the Job Corps illustration, the treatment is program assignment; the long-term outcomes are 4-year weekly earnings and weeks employed; the surrogates are years 2–3 earnings and employment; and the proxies are GED possession SRdS\in\mathbb R^d01 and English mother tongue SRdS\in\mathbb R^d02, together with baseline SRdS\in\mathbb R^d03. The reported finding is that naive surrogate-index methods underestimate the ATE, even when proxies are simply added as regressors, whereas the proximal estimator recovers the randomized benchmark, though with wider standard errors (Hung et al., 25 Jan 2026).

A complementary line of work treats surrogacy as a sensitivity parameter rather than a maintained assumption. The Weighted Surrogate Index for a known copula SRdS\in\mathbb R^d04 is

SRdS\in\mathbb R^d05

where SRdS\in\mathbb R^d06. When SRdS\in\mathbb R^d07 is the independence copula, SRdS\in\mathbb R^d08 and SRdS\in\mathbb R^d09, so the classical surrogate index is recovered. If the copula is known, the ATE on SRdS\in\mathbb R^d10 is identified by the ATE on the WSI; if the copula is unknown but bounded in concordance order between SRdS\in\mathbb R^d11 and SRdS\in\mathbb R^d12, then the sharp identified set is SRdS\in\mathbb R^d13. The Fréchet–Hoeffding bounds yield worst-case intervals, with the upper bound expressed through conditional SRdS\in\mathbb R^d14 functionals (Fan et al., 28 Feb 2026).

This sensitivity framework also develops debiased, Neyman-orthogonal moments and cross-fitted estimators for both point-identified and partially identified cases. In the Pakistan poverty-alleviation application, even small departures from surrogacy, with SRdS\in\mathbb R^d15, can render the long-term effect statistically significant and change its sign. The report further notes that the tight centering of SRdS\in\mathbb R^d16 makes the results relatively insensitive to the copula family, so the dominant feature is the global strength of dependence summarized by Kendall’s tau (Fan et al., 28 Feb 2026).

5. Trial-level surrogate indices in high-dimensional meta-analysis

In multi-trial surrogate evaluation, the object of interest is not the individual-level imputation map SRdS\in\mathbb R^d17 but trial-level surrogacy across studies. RISE-Meta addresses this problem when the candidate surrogates are high-dimensional, such as omics data. For each candidate marker SRdS\in\mathbb R^d18 in trial SRdS\in\mathbb R^d19, it computes a within-study nonparametric surrogacy estimator based on the indicator kernel

SRdS\in\mathbb R^d20

Using treated–control pairs, it estimates the “area under the Mann–Whitney curve” for the primary endpoint and for the marker, denoted SRdS\in\mathbb R^d21 and SRdS\in\mathbb R^d22, and defines the trial-level bias

SRdS\in\mathbb R^d23

If SRdS\in\mathbb R^d24, the average superiority probability on the surrogate equals that on the primary endpoint, so the candidate is an unbiased trial-level surrogate (Hughes et al., 5 May 2026).

Across trials, RISE-Meta fits the hierarchical Gaussian model

SRdS\in\mathbb R^d25

with SRdS\in\mathbb R^d26 estimated by REML and inverse-variance weights SRdS\in\mathbb R^d27. The pooled surrogate effect is

SRdS\in\mathbb R^d28

and uncertainty is quantified by the Hartung–Knapp–Sidik–Jonkman variance estimator, yielding a SRdS\in\mathbb R^d29 reference distribution (Hughes et al., 5 May 2026).

Operational validity is defined by equivalence, not by exact equality. For a pre-specified equivalence bound SRdS\in\mathbb R^d30, trial-level surrogacy is declared when SRdS\in\mathbb R^d31, tested through TOST: SRdS\in\mathbb R^d32 Equivalence is declared at level SRdS\in\mathbb R^d33 when the resulting SRdS\in\mathbb R^d34-value satisfies SRdS\in\mathbb R^d35, equivalently when the SRdS\in\mathbb R^d36 confidence interval lies inside SRdS\in\mathbb R^d37. When the number of markers SRdS\in\mathbb R^d38 is large, SRdS\in\mathbb R^d39-values are adjusted by Benjamini–Hochberg to give SRdS\in\mathbb R^d40 (Hughes et al., 5 May 2026).

The framework then forms a composite surrogate signature from the promising markers SRdS\in\mathbb R^d41. After standardizing each selected marker to SRdS\in\mathbb R^d42, it defines

SRdS\in\mathbb R^d43

and the weights

SRdS\in\mathbb R^d44

The composite index is

SRdS\in\mathbb R^d45

The paper states that, by construction, SRdS\in\mathbb R^d46 concentrates information in markers with small SRdS\in\mathbb R^d47 and low variance, improving overall surrogacy through SRdS\in\mathbb R^d48 closer to zero, narrower confidence and prediction intervals, and higher concordance (Hughes et al., 5 May 2026).

The stated assumptions include independence of trials, i.i.d. sampling within trials, fully nonparametric within-study estimation, normality only for study-level SRdS\in\mathbb R^d49, and clinically or power-motivated selection of SRdS\in\mathbb R^d50. The paper notes that TOST is conservative when SRdS\in\mathbb R^d51 is small, that Bayesian shrinkage for SRdS\in\mathbb R^d52 could improve power, and that avoiding the surrogate paradox requires monotonicity and no residual treatment effect conditions. The method is reported to be directly applicable to both high-dimensional and classical low-dimensional meta-analyses, with agreement with the Buyse et al. approach when SRdS\in\mathbb R^d53 is small; applications include gene expression as trial-level surrogate markers for antibody response to seasonal influenza vaccination (Hughes et al., 5 May 2026).

6. Other technical meanings and terminological distinctions

Outside causal-inference applications, the phrase “surrogate index” appears in surrogate modeling with a different meaning. In the Additive Multi-Index Gaussian process model, a surrogate index is a low-dimensional linear combination of a high-dimensional input SRdS\in\mathbb R^d54,

SRdS\in\mathbb R^d55

that captures one “mode” or “physics” of variation in the response. The surrogate model is additive,

SRdS\in\mathbb R^d56

with independent GP priors on the SRdS\in\mathbb R^d57, Laplace priors on entries of SRdS\in\mathbb R^d58, and a variational approximation over inducing variables and index matrices. The stated purpose is to reduce dimensionality, avoid the curse of dimensionality of a full SRdS\in\mathbb R^d59-dimensional GP, and separate distinct physical mechanisms. In the quark–gluon plasma application, the learned indices are reported to align with known physics, with one index selecting only the three initial-condition inputs and another selecting only viscosity-related parameters (Li et al., 2023).

In Multi-Index Stochastic Collocation, the surrogate index is not a scalar prediction map at all. It is an admissible downward-closed set SRdS\in\mathbb R^d60 of multi-indices SRdS\in\mathbb R^d61 that simultaneously tracks solver-fidelity levels and polynomial or interpolation levels. The MISC surrogate is

SRdS\in\mathbb R^d62

and an adaptive profit-driven algorithm expands SRdS\in\mathbb R^d63 while enforcing admissibility. The PlateauMISC refinement monitors the spectral envelope SRdS\in\mathbb R^d64, detects solver-noise plateaux by a piecewise-linear fit in SRdS\in\mathbb R^d65, and blocks refinements whose contributions lie entirely beyond the detected plateau degree. This use of “surrogate index” refers to the combinatorial scaffold of a multi-fidelity surrogate, not to a surrogate endpoint or a causal bridge (Kent et al., 4 Jul 2025).

These usages are structurally related only at a high level. In each case, the index is a summary object that organizes surrogate information, but the inferential target changes: long-term treatment effects in the classical and proximal formulations, trial-level surrogacy across studies in RISE-Meta, low-dimensional latent structure in additive GP surrogates, and admissible refinement structure in MISC. A plausible implication is that the term should always be interpreted relative to its surrounding methodological framework rather than as a domain-independent object.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Surrogate Index.