---
title: Nonparametric Bridge Functions
url: https://www.emergentmind.com/topics/nonparametric-bridge-functions
type: topic
---

# Nonparametric Bridge Functions

Searching arXiv for the cited papers to ground the article in the current literature.
arXiv search: 2604.08681, 2605.20007, 2009.01059, 2108.09574, 1405.5929
Nonparametric bridge functions are functions defined through conditional expectation or integral-equation restrictions that connect observed measurements or proxies to a target latent or interventional quantity without imposing a parametric model on the bridge itself. In recent causal-inference work, they appear in two closely related settings. In latent-outcome problems, a bridge function maps each imperfect indicator onto a common benchmark scale in expectation conditional on the latent variable, thereby supporting identification of the average latent treatment effect. In proximal causal inference, bridge functions solve Fredholm equations of the first kind that recover interventional distributions from proxy variables in the presence of latent confounding, and extended bridge functions enlarge this framework to joint interventional distributions that retain the proxies themselves [2604.08681] [2605.20007].

## 1. Conceptual role and problem setting

A nonparametric bridge function is introduced when the object of interest is not directly observed or is confounded by latent structure, while the available data contain multiple noisy indicators or proxy variables. In the latent-outcome setting, the target is the average latent treatment effect
\[
\tau = \mathbb{E}[\eta_i(1)-\eta_i(0)],
\]
where the latent outcome \(\eta_i\) is observed only through imperfect measurements \(Y_{ij}\). The central difficulty is that different indicators may have different and possibly nonlinear relationships with the same latent outcome, so raw measurements are not directly comparable within or across studies [2604.08681].

The same bridge-function logic appears in proximal causal inference, but with a different target. There, the problem is latent confounding by \(U\) in the presence of an outcome-inducing proxy \(W\), a treatment-inducing proxy \(Z\), treatment \(A\), and covariates \(X\). Standard outcome and treatment bridge functions are introduced to identify \(\dodistr{Y}{}{a}\) without observing \(U\), and extended bridge functions are introduced because standard bridges do not generally identify joint interventional distributions that contain all proxy variables used to define them [2605.20007].

This suggests a unifying view: nonparametric bridge functions are inverse-problem objects that “translate” observed proxy information into a target scale or target distribution. In the latent-outcome case the target is a benchmark measurement scale; in proximal inference the target is an interventional distribution.

## 2. Measurement bridge functions for latent outcomes

In the latent-outcome framework, a benchmark measurement \(Y_{i1}\) is selected and assumed to satisfy
\[
\mathbb{E}[Y_{i1}\mid \eta_i] = \eta_i.
\]
This centering measurement anchors the latent construct on a usable scale. The other indicators need not be linear or unbiased; they may be nonlinear, noisy, or differently coded [2604.08681].

For each auxiliary measurement \(Y_{ij}\), \(j=2,\dots,J\), the measurement bridge function \(\phi_j\) is defined by
\[
\mathbb{E}[Y_{i1}\mid \eta_i] = \mathbb{E}[\phi_j(Y_{ij})\mid \eta_i].
\]
The defining requirement is therefore equality in conditional expectation given the latent outcome, not equality of the raw measurements. Once such a \(\phi_j\) is obtained, the transformed quantity \(\phi_j(Y_{ij})\) is interpretable on the benchmark scale [2604.08681].

The transformed measurements can then be aggregated as
\[
\tilde Y_i = \sum_{j=1}^J \omega_j \phi_j(Y_{ij}), \qquad \sum_{j=1}^J \omega_j = 1,
\]
often using inverse-variance weights for efficiency. Because the benchmark measurement is centered, the average latent treatment effect can be expressed through transformed measurements as a linear functional of the bridge function. For \(j \ge 2\),
\[
\tau = \mathbb{E}[s(Z_i)\phi_j(Y_{ij})],
\qquad
s(Z_i)=\frac{Z_i}{\pi} - \frac{1-Z_i}{1-\pi},
\qquad
\pi=\Pr(Z_i=1).
\]
This linear-functional structure is central for estimation under weak first-stage identification [2604.08681].

The framework is explicitly motivated by two noncomparability problems. The first is study noncomparability: different studies may use different indicator sets, so standard summary-index methods can target different empirical objects even when the underlying latent treatment effect is the same. The second is measurement noncomparability within a study: different indicators may relate to \(\eta_i\) in different, possibly nonlinear ways. The bridge-function construction addresses both by mapping every indicator to a common benchmark in expectation [2604.08681].

## 3. Existence and identification as nonparametric inverse problems

The measurement bridge equation is written as
\[
K\phi = r,
\]
where \(K\phi = \mathbb{E}[\phi(Y_{ij})\mid \eta_i]\) and \(r(\eta_i)=\mathbb{E}[Y_{i1}\mid \eta_i]\). The paper characterizes this as a Fredholm integral equation of the first kind. A key sufficient condition for existence is completeness:
\[
\mathbb{E}[g(\eta_i)\mid Y_{ij}] = 0 \text{ a.s.} \iff g(\eta_i)=0 \text{ a.s.}
\]
for every square-integrable function \(g\). Under regularity conditions and completeness, Proposition 1 establishes existence of a bridge function \(\phi_j\) satisfying the defining conditional expectation restriction [2604.08681].

Identification is formulated as a nonparametric instrumental variables problem. If there exists an auxiliary variable \(W_i\) such that the required mean-independence restrictions hold and the instrument is complete in the sense that
\[
\mathbb{E}[g(Y_{ij})\mid W_i] = 0 \text{ a.s.} \implies g(Y_{ij})=0 \text{ a.s.},
\]
then \(\phi_j\) is uniquely identified from
\[
\mathbb{E}[\phi_j(Y_{ij})\mid W_i] = \mathbb{E}[Y_{i1}\mid W_i].
\]
The paper notes that valid instruments can include treatment assignment \(Z_i\), other measurements, and pre-treatment covariates \(X_i\) [2604.08681].

In proximal causal inference, the same inverse-problem structure appears in bridge equations involving the proxies \(W\) and \(Z\). The standard outcome bridge \(h(y,w,a,x)\) solves
\[
\sum_w h(y,w,a,x)\,p(w\mid z,a,x)=p(y\mid z,a,x),
\]
and the standard treatment bridge \(q(z,a,x)\) solves
\[
\sum_z q(z,a,x)\,p(z\mid w,a,x)=\frac{1}{p(a\mid w,x)}.
\]
The corresponding completeness conditions are stated as conditional completeness of \(Z\) for \(U\) given \((A,X)\) for outcome bridges and of \(W\) for \(U\) given \((A,X)\) for treatment bridges [2605.20007].

A common misconception is that “nonparametric” eliminates structural assumptions. The opposite is true here: the framework replaces parametric restrictions on the bridge function with operator-level assumptions such as completeness, mean independence, proxy validity, positivity, and solvability of Fredholm equations.

## 4. Estimation, debiasing, and weak identification

Estimation of measurement bridge functions proceeds in two stages. In the first stage, the bridge is estimated by NPIV through a penalized minimax problem with cross-fitting,
\[
\hat\phi_j^{(-k)} = \arg\min_{\phi\in \Phi_n} \sup_{q\in Q_n} \mathbb{E}_{n,-k}\!\left[ \phi(Y_{ij}) - Y_{i1} + q(W_i) \right]^2 - \gamma_n\|q\|_{Q}^2 + \mu_n\|\phi\|_{\Phi}^2.
\]
Because the first stage is an ill-posed inverse problem, regularization is necessary [2604.08681].

The second stage uses a Neyman orthogonal score,
\[
\psi_{ij} = \alpha_{0j}(R_i)\,\phi_j(Y_{ij}) + q_{0j}(W_i)\bigl(Y_{i1}-\phi_j(Y_{ij})\bigr) - \tau,
\]
where \(\alpha_{0j}(R_i)\) is a Riesz representer and \(q_{0j}(W_i)=\mathbb{E}[\xi_{0j}(Y_{ij})\mid W_i]\) is a debiasing nuisance. The orthogonality construction removes first-order sensitivity to estimation error in the first-stage bridge estimator [2604.08681].

A major theoretical point is that the bridge function itself may be weakly identified while the causal estimand remains strongly identified. The paper emphasizes that the average latent treatment effect is a smooth linear functional of the bridge function, so valid root-\(n\) inference remains possible even when the nuisance bridge is unstable. Under the stated conditions,
\[
\sqrt{n}(\hat\tau-\tau)\overset{d}{\to} N(0,\Omega).
\]
The same logic applies to downstream regression coefficients obtained from bridge-transformed outcomes [2604.08681].

The simulation evidence is designed to compare cross-study comparability when the true latent effect is identical but measurement systems differ. The reported results are as follows.

| Method | Cross-study gap | Rejection rate |
|---|---:|---:|
| PCA | \(0.256\) | \(0.247\) |
| ICW | \(0.366\) | \(1.000\) |
| WSI | \(0.072\) | \(0.012\) |
| NSI | \(0.004\) | \(0.006\) |

These results are used to argue that PCA and inverse covariance weighting can generate spurious cross-study differences, whereas the nonparametric scaled index restores comparability. The same paper also reports that nonparametric estimators can be less stable and need larger samples, so when the measurement relationship is plausibly linear, the linear WSI approach may be preferable for efficiency [2604.08681].

## 5. Extended bridge functions and joint interventional distributions

Extended bridge functions generalize standard proximal bridge functions by retaining an additional proxy variable in the target object. The extended outcome bridge \(h(y,w,w',a,x)\) solves
\[
\sum_{w'} h(y,w,w',a,x)\,p(w'\mid z,a,x)=p(y,w\mid z,a,x),
\]
and the extended treatment bridge \(q(z,z',a,x)\) solves
\[
\sum_{z'} q(z,z',a,x)\,p(z'\mid w,a,x)=\frac{p(z\mid w,x)}{p(a\mid w,x)}.
\]
Standard bridges are recovered by marginalization:
\[
\tilde{h}(y,w,a,x)=\sum_{w'} h(y,w',w,a,x), \qquad
\tilde{q}(z,a,x)=\sum_{z'} q(z',z,a,x).
\]
Extended bridges are therefore a strict generalization of the standard bridge formalism [2605.20007].

Under stronger assumptions, extended bridge functions identify joint interventional distributions containing the proxies:
\[
\dodistr{y,w,z,x}{}{a} = \sum_{w'} h(y,w,w',a,x)\,p(w',z,x),
\]
and
\[
\dodistr{y,w,z,x}{}{a} = \sum_{z'} q(z,z',a,x)\,p(y,w,z',a,x).
\]
The importance of these results is that many intermediate objects in generalized proximal identification algorithms are kernels that still contain proxies. Standard bridge functions generally identify only marginals such as \(\dodistr{Y}{}{a}\), or at most distributions excluding one of the proxies, whereas extended bridges identify joint targets that retain both proxies [2605.20007].

The paper reformulates these results in operator language. Conditional expectation operators \(T_{a,x}\) and \(T'_{a,x}\) map functions of \(W\) and \(Z\) into conditional expectations, and existence of bridge functions is tied to completeness and regularity of these operators via spectral decompositions. The generalized proximal identification algorithm then uses the kernel operations \(\{Fix,Obf,Tbf,Ebf\}\), together with district factorization, to identify \(\dodistr{Y}{}{a}\) when an appropriate sequence of kernel operations exists [2605.20007].

This suggests a broader significance for nonparametric bridge functions beyond a single estimating equation: they serve as modular identification objects inside larger kernel-based causal algorithms.

## 6. Assumptions, design implications, and terminological boundaries

The practical guidance in the latent-outcome framework is explicit. A common benchmark measurement should be included across studies whenever cross-study comparison is desired. Indicators that discard information should be avoided if possible, because completeness matters for identification. Measurements that are informative and monotonic in the latent outcome are preferred, and measurement design should be treated as part of experimental design rather than as an afterthought [2604.08681].

In proximal inference, the bridge equations require both causal and nonparametric assumptions. The causal side includes latent ignorability, proxy validity, and positivity; the nonparametric side includes completeness and operator solvability. For the extended bridge results, stronger assumptions are used, including exclusion restrictions \(W(a)=W\), \(Z(a)=Z\), \(X(a)=X\) and ignorability given \(W,Z,U,X\) [2605.20007].

A second misconception is that bridge functions are merely ad hoc reweighting devices. In the measurement setting they are benchmark-scale mappings defined by conditional expectation restrictions; in proximal inference they are solutions to exact integral equations whose identifying role is tied to latent-variable structure. Their inferential validity depends on the bridge equations and associated assumptions, not simply on predictive fit.

The term “bridge function” also has an unrelated meaning in liquid-state theory. There, the bridge function is the formally exact residual term in the closure
\[
g(r)=\exp\!\left[-\beta u(r)+h(r)-c(r)+B(r)\right],
\]
with \(B(r)\) representing the part omitted by the hypernetted-chain approximation. Recent work in that literature studies simulation-based extraction, parametrization, and approximate invariance of \(B(r)\) for Yukawa and one-component plasmas, as well as bridge corrections in classical DFT plus RISM for liquid water [2009.01059] [2108.09574] [1405.5929]. This is a separate research tradition from the causal-inference usage of nonparametric bridge functions, despite the shared terminology.

Within causal inference, the current literature presents nonparametric bridge functions as a general strategy for problems with latent outcomes or latent confounding: identify a bridge through conditional moment restrictions, regularize the resulting inverse problem, and target estimands that remain well behaved even when the bridge itself is weakly identified.

Source: https://www.emergentmind.com/topics/nonparametric-bridge-functions