---
title: Neyman-Orthogonal Moments
url: https://www.emergentmind.com/topics/neyman-orthogonal-moments-831913cc-c9f2-437b-a73b-6a178d592c63
type: topic
---

# Neyman-Orthogonal Moments

Neyman-orthogonal moments are moment or score functions constructed so that small perturbations of nuisance parameters do not change the first-order expectation of the estimating equation at the truth. In semiparametric inference and debiased or double machine learning, this property is used to suppress first-stage bias from high-dimensional, nonparametric, or machine-learning nuisance estimators while retaining $\sqrt{n}$-scale inference for low-dimensional targets. The modern literature treats orthogonal moments as a unifying device linking semiparametric efficiency, influence-function constructions, Riesz representers, balancing methods, higher-order robustness, and extensions to Bayesian and survival settings [2605.06386] [2303.11418] [1711.00342].

## 1. Definition and canonical construction

A standard semiparametric setup posits observable data $Z$, a finite-dimensional parameter of interest $\theta$, an infinite-dimensional nuisance $\eta$, and a moment function
\[
g(Z;\theta,\eta)\in\mathbb R^k
\]
such that
\[
\E_{P_{\theta_0,\eta_0}}\bigl[g(Z;\theta_0,\eta_0)\bigr]=0.
\]
Neyman orthogonality is the requirement that, for every nuisance direction $h$, the Gateaux derivative of the population moment with respect to nuisance perturbations vanishes at the truth:
\[
\left.\frac{d}{dt}\right|_{t=0}\;
\E_{P_{\theta_0,\eta_0+t\,h}}
\bigl[g(Z;\theta_0,\eta_0+t\,h)\bigr]=0.
\]
Equivalently, when differentiation under the integral sign is valid and $S_\eta(Z)[h]$ denotes the nuisance score, orthogonality is the condition
\[
\E\bigl[g(Z;\theta_0,\eta_0)\,S_\eta(Z)[h]\bigr]=0
\quad \forall\,S_\eta(\cdot)[h]\in T_{\eta_0},
\]
with $T_{\eta_0}$ the nuisance tangent space [2303.11418].

A particularly transparent construction arises when the target is a linear functional of a regression function. Let $W=(X,Y)$, let $\gamma_0(x)=E[Y\mid X=x]$, and suppose
\[
\theta_0 = E\big[m(W;\gamma_0)\big],
\]
with $m$ linear in $\gamma_0$. By the Riesz representation theorem in $L_2(P_X)$ there exists $\alpha_0\in L_2(P_X)$ such that
\[
E\big[m(W;\gamma)\big] = E\big[\alpha_0(X)\gamma(X)\big]
\]
for every $\gamma\in L_2(P_X)$. The associated orthogonal score is
\[
\psi\big(W;\eta_0,\theta_0\big)
=
\alpha_0(X)\{Y-\gamma_0(X)\}
+
m(W;\gamma_0)
-
\theta_0,
\qquad
\eta_0=(\alpha_0,\gamma_0).
\]
By construction,
\[
E\big[\psi(W;\eta_0,\theta_0)\big]=0,
\]
and $\theta_0$ solves $E[\psi(W;\eta_0,\theta)]=0$ [2605.06386].

This formulation places orthogonal moments at the intersection of influence-function theory and linear functional estimation. The nuisance pair $(\alpha_0,\gamma_0)$ is not ancillary to the score; it is the mechanism through which the score neutralizes first-order nuisance perturbations.

## 2. Gateaux orthogonality and first-stage insensitivity

The defining operational feature of a Neyman-orthogonal moment is first-order insensitivity to nuisance estimation error. In the Riesz-based construction above, for any perturbation direction $\delta\eta=(\delta\alpha,\delta\gamma)$,
\[
\left.\frac{d}{dt}\;
E\Big[\psi\big(W;\eta_0+t\,\delta\eta,\theta_0\big)\Big]\right|_{t=0}
=0.
\]
The cancellation follows from two facts recorded in the literature: first, $Y-\gamma_0(X)$ is mean-zero under the model; second, the pathwise derivative of the linear functional is represented by the Riesz representer, so the derivative contributions in $\delta\gamma$ offset each other exactly [2605.06386].

This property underlies the usual plug-in equation
\[
\frac1n\sum_{i=1}^n s(O_i;\hat\theta,\hat\eta)=0.
\]
If $s$ is orthogonal, a Taylor expansion in $\eta$ shows that
\[
E_0[s(O;\theta_0,\hat\eta)]-E_0[s(O;\theta_0,\eta_0)]
\]
is of second order in $\hat\eta-\eta_0$, so even a relatively slow $n^{-1/4}$-rate estimate of $\eta_0$ perturbs the moment by $o(n^{-1/2})$. This is the standard first-order robustness statement behind double or debiased machine learning [2602.20371]. In the formulation of higher-order orthogonality, first-order orthogonality corresponds to the case $k=1$, and the familiar $o(n^{-1/4})$ requirement appears as the special case of the general rate $o(n^{-1/(2k+2)})$ [1711.00342].

Cross-fitting is repeatedly used to operationalize this insensitivity. In the high-dimensional and higher-order developments, cross-fitting separates nuisance training from score evaluation so that the leading empirical process is driven by the orthogonal score at the truth rather than by overfitting artifacts. In the terminology of the higher-order theory, without cross-fitting one pays a first-stage bias penalty that can destroy the $\sqrt n$ rate [1711.00342].

## 3. ATE scores, Riesz representers, and balancing

A canonical example is the average treatment effect under binary treatment. In the notation $O=(Y,Z,X)$ with $Z\in\{0,1\}$, define
\[
\mu_0(z,x)=E[Y\mid Z=z,X=x],
\qquad
e_0(x)=P(Z=1\mid X=x).
\]
The ATE
\[
\theta_0=E[\mu_0(1,X)-\mu_0(0,X)]
\]
solves
\[
E_0[\psi(O;\theta,\mu,e)]=0,
\]
where the efficient influence function is
\[
\psi(O;\theta,\mu,e)
=
\{\mu(1,X)-\mu(0,X)-\theta\}
+
\frac{Z\,[Y-\mu(1,X)]}{e(X)}
-
\frac{(1-Z)\,[Y-\mu(0,X)]}{1-e(X)}.
\]
At $(\theta_0,\mu_0,e_0)$, the derivative in the directions $\delta\mu,\delta e$ is zero, so the AIPW score is Neyman-orthogonal [2602.20371].

The recent balancing literature reframes this construction in terms of the error term that enters the orthogonal score. In the heterogeneous ATE problem with regressor $X=(D,Z)$ and outcome regression $\gamma_0(d,z)=E[Y\mid D=d,Z=z]$, the deterministic part of the error of the sample-moment equation can be written in terms of the regression error $\xi(d,z)=\gamma_0(d,z)-\widehat\gamma(d,z)$ as
\[
\frac{1}{n}\sum_i
\Bigl\{
\widehat\alpha(D_i,Z_i)\,\xi(D_i,Z_i)
-
\bigl[\xi(1,Z_i)-\xi(0,Z_i)\bigr]
\Bigr\}.
\]
When $\xi(d,z)$ lies in the span of a finite set of covariate-only functions $\widetilde\Phi_1(z),\ldots,\widetilde\Phi_p(z)$, covariate-balancing constraints force the deterministic bias to vanish on that subspace. This is exactly how entropy- or kernel-based covariate balancing methods work when the relevant regression error depends only on $Z$ [2605.06386].

The position advanced in recent work is more specific than a general endorsement of balancing. The argument is that, in debiased machine learning, balancing functions should be derived from the Neyman orthogonal score, not chosen only as functions of covariates. For ATE estimation under treatment effect heterogeneity, the score error generally contains treatment-specific components because the outcome regression is a function of the full regressor $X=(D,Z)$. In that case, balancing common functions of $Z$ can leave the treatment-specific component unbalanced. The proposed general principle is therefore regressor balancing, implemented by Riesz regression with basis functions of $X$:
\[
\widehat\beta
=
\arg\min_\beta
\Bigl\{
\frac1n\sum_{i=1}^n\alpha_\beta(X_i)^2
-
\frac{2}{n}\sum_{i=1}^n m\bigl(W_i;\alpha_\beta\bigr)
+
\lambda\,J(\beta)
\Bigr\},
\qquad
\alpha_\beta(X)=\beta^\top\Phi(X).
\]
Its first-order condition implies
\[
\Delta_n\bigl(\widehat\alpha,\Phi_j\bigr)
=
-\tfrac{\lambda}{2}\,\partial_jJ(\widehat\beta),
\]
so in the unregularized case $(\lambda=0)$ it exactly balances each $\Phi_j$. Covariate balancing is therefore not treated as invalid; it is treated as the special case appropriate when the score-relevant regression error is a function of covariates alone [2605.06386].

## 4. Existence, restricted local non-surjectivity, and information

The existence of orthogonal moments is not automatic in general semiparametric models. The central criterion introduced in recent theory is Restricted Local Non-surjectivity (RLN). Under the null $\theta=\theta_0$, let
\[
T_0
=
\bigl\{S_\eta(\cdot)[h]: h\in\dot{\mathcal H},\ \text{paths with }\theta_t=\theta_0\bigr\}
\]
be the restricted nuisance tangent space, with closure $\overline T_0$ in $L^2_0$. The model satisfies Restricted Local Surjectivity if $\overline T_0=L^2_0$, and otherwise satisfies RLN. In operator language,
\[
\mathrm{Range}\bigl(S_{\rm nuis}\bigr)\subsetneq L^2_0(P_0)
\]
is the RLN condition [2303.11418].

The main existence theorem states that, under standard regularity, the model satisfies RLN if and only if there exists a nonzero orthogonal moment:
\[
g\in G_0\cap[T_0]^\perp.
\]
Equivalently,
\[
\Bigl\{g\in G_0:\E[g\,S_\eta(\cdot)[h]]=0\ \forall h\Bigr\}
\text{ is nontrivial}
\iff
\overline T_0\subsetneq L^2_0.
\]
An adjoint formulation gives the same criterion as
\[
\ker(S_\eta^*)\neq\{0\}.
\]
This result separates the existence question from identification of the parameter of interest and from identification of the nuisance parameter: RLN does not require either one [2303.11418].

Existence, however, is not the same as informativeness. Let
\[
S_{\rm eff}=S_\theta-\Pi_{T_0}S_\theta
\]
be the efficient score for $\theta$, and let
\[
I_{\theta\theta}
=
\E\bigl[S_{\rm eff}S_{\rm eff}^\prime\bigr].
\]
An orthogonal moment has nontrivial local power if and only if $S_{\rm eff}\not\equiv0$, equivalently if and only if the efficient Fisher Information matrix is non-zero, though possibly singular. In the scalar case, “nonzero” and “full-rank” coincide; in the multivariate case, a singular information matrix corresponds to semi-identification along certain directions [2303.11418].

This distinction addresses a recurrent misconception. Orthogonal moments can exist in abundance, yet still fail to be informative for the parameter of interest if the efficient score vanishes. Conversely, the existence criterion explains why orthogonal moments can be found in models with unobserved heterogeneity, conditional moment restrictions with possibly different conditioning variables, fully saturated two stage least squares, heterogeneous parameters in treatment effects, sample selection models, and popular models of demand for differentiated products [2303.11418].

## 5. Higher-order orthogonality and robustness limits

First-order Neyman orthogonality can be generalized to $k$-th order orthogonality. Let $h_0(X)$ be the nuisance and let $\alpha$ be a multi-index. A moment $m$ is $k$-orthogonal at $(\theta_0,h_0)$ if
\[
E\bigl[D^\alpha m\,(Z,\theta_0,h_0(X))\mid X\bigr]=0
\quad\text{for every } |\alpha|\le k.
\]
A Taylor expansion of the plug-in moment around $h_0(X)$ then removes all nuisance terms up to order $k$, leaving only the $(k+1)$-st order remainder. Under standard regularity and cross-fitting, if $\hat h$ estimates $h_0$ at rate
\[
o\bigl(n^{-1/(2k+2)}\bigr)
\quad\text{in }L^{2(k+1)},
\]
then
\[
\sqrt n\,(\hat\theta-\theta_0)\xrightarrow{d}N\bigl(0,J^{-1}VJ^{-1}\bigr).
\]
For $k=1$ this recovers the $o(n^{-1/4})$ first-order result; for $k=2$ the required nuisance rate becomes $o(n^{-1/6})$ [1711.00342].

The partially linear regression model is the principal explicit case:
\[
Y=\theta_0\,T+f_0(X)+\epsilon,
\qquad
E[\epsilon\mid X,T]=0,
\]
\[
T=g_0(X)+\eta,
\qquad
E[\eta\mid X]=0.
\]
In this setting, second-order orthogonal moments can be constructed if and only if the treatment residual is not normally distributed. The characterization is
\[
E[\eta^{r+1}\mid X]
=
r\,E[\eta^2\mid X]\,E[\eta^{r-1}\mid X]
\quad \forall r\ge2
\]
if and only if the conditional distribution of $\eta\mid X$ is Gaussian almost surely. Hence a violation such as non-zero skewness or excess kurtosis permits second-order orthogonal moments; conditional Gaussianity implies that no non-degenerate second-order orthogonal moment exists [1711.00342].

This yields a precise limitation rather than a generic promise of arbitrary robustness. Higher-order orthogonality can improve the admissible nuisance rate, but it requires deeper moment constructions and stronger smoothness. The same source states that in high-dimensional linear nuisances, second-order orthogonality can tolerate sparsity up to $s=o(n^{2/3}/\log p)$ versus $s=o(n^{1/2}/\log p)$ for first-order, while in nonparametric settings the needed $L^{2(k+1)}$ rate may require very strong smoothness [1711.00342].

## 6. Bayesian and high-dimensional survival extensions

Neyman orthogonality has also been exported beyond the conventional frequentist DML setting. In semi-parametric Bayesian inference with a non-parametrically modelled nuisance component, the relevant result is that when the nuisance and targeted parameters satisfy a Neyman orthogonal score property, cutting feedback through a two-step procedure is a valid way of conducting Bayesian inference. Using a Dirichlet process and the Bayesian bootstrap, one obtains that the marginal posterior of the targeted parameter exhibits good frequentist properties despite not accounting for the inferential uncertainty of the nuisance parameter. For the plug-in estimator $\hat\theta$ and the Bayesian-bootstrap draw $\hat\theta_{BB}$,
\[
\sqrt n(\hat\theta-\theta_0)\to_{\mathcal D} N(0,\Sigma),
\qquad
\sqrt n(\hat\theta_{BB}-\hat\theta)\mid O_1^n \to_{\mathcal D} N(0,\Sigma)
\]
with the same asymptotic variance. The same work also investigates the absence of Neyman orthogonality and shows that, for a simple family of useful scores, the posterior distribution is asymptotically unchanged by nuisance estimation provided the nuisance estimator is consistent [2602.20371].

In survival analysis, orthogonality is used to debias high-dimensional hazard-based estimators. The HSCI framework observes
\[
\mathcal O_i=(X_i,\delta_i,D_i,Z_i),
\]
assumes unconfoundedness and independent censoring given $(D,Z)$, and posits a sparse high-dimensional Cox proportional hazards outcome model together with a high-dimensional logistic propensity score working model. The Neyman near-orthogonal score is
\[
\Psi(\{\mathcal O_i\};\theta,\eta)
=
\dot\ell_\theta(\theta,\beta)-\mu^\top\dot\ell_\beta(\theta,\beta),
\qquad
\eta=(\beta,\mu),
\]
with
\[
\mu_{a0}=\Sigma_{\beta\beta}^{-1}\Sigma_{\beta\theta}
\]
chosen so that
\[
\partial_\eta\EE[\Psi(\theta_0,\eta_{a0})]=0
\quad\text{(near-orthogonality)}.
\]
Implemented with cross-fitting, the framework establishes root-n asymptotic normality and consistent variance estimation under doubly robust nuisance-rate conditions:
\[
\sqrt n\,(\hat\theta-\theta_0)\dto N(0,\sigma^2),
\qquad
\sigma^2
=
J_0^{-2}\,\EE\Bigl[n\,\Psi(\theta_0,\eta_{a0})^2\Bigr].
\]
A Wald-type $95\%$ confidence interval is
\[
\hat\theta\pm1.96\,\hat\sigma/\sqrt n.
\]
The same framework extends to inference on high-dimensional survival covariate effects by one-step correction of the Lasso estimator [2606.14132].

Across these extensions, the recurring pattern is stable. Orthogonality is used to decouple the target parameter from first-order nuisance distortion, but the technical realization varies with the inferential regime: efficient influence functions in causal inference, RLN and tangent-space geometry in existence theory, higher-order derivatives in robustness analysis, Dirichlet-process and Bayesian-bootstrap arguments in Bayesian semiparametrics, and near-orthogonal score correction in high-dimensional Cox models. A plausible implication is that Neyman-orthogonal moments are best viewed not as a single estimator, but as a design principle for constructing inferentially stable estimating equations in semiparametric problems.

Source: https://www.emergentmind.com/topics/neyman-orthogonal-moments-831913cc-c9f2-437b-a73b-6a178d592c63