- The paper introduces a unified nearest neighbor graph-based estimator for testing conditional mean independence and estimating Sobol' indices, ensuring computational efficiency.
- It establishes an asymptotically normal test statistic that enables model-free variable screening and offers tight control over type I error in complex settings.
- The methodology extends to higher-order Sobol' indices to quantify interaction effects, bridging nonparametric regression diagnostics and global sensitivity analysis.
Nearest Neighbor Graph Estimation of Conditional Mean Independence and Sobol' Indices
Introduction and Motivation
Quantification of the influence of covariates on a response via their effects on the conditional mean function is central to regression, model evaluation, variable selection, and global sensitivity analysis (GSA). In "Conditional Mean Independence and Global Sensitivity Analysis using Nearest Neighbor Graphs" (2607.04692), the authors introduce a unified, computationally efficient, nonparametric approach for testing conditional mean independence, estimating variable importance (via variance-based indices), and conducting GSA, all grounded in nearest neighbor (NN) graph methodology.
The approach is built on a normalized, variance-based measure of mean dependence, closely related to classical Sobol' indices. The authors propose a consistent, rate-optimal NN graph estimator for this measure that is computation-efficient and scalable to high dimensions. The framework yields an asymptotically normal statistic for conditional mean independence, facilitating universally consistent, non-splitting hypothesis testing and model-free variable screening. The methodology is further extended to estimation of higher-order (interaction) Sobol' indices.
Conditional Mean Dependence and Sobol' Indices
Let (Y,X) be jointly distributed, with Y∈Rp, X∈Rd. The authors focus on the normalized conditional mean discrepancy (NCMD):
η=E[∥Y−E[Y]∥22​]E[∥E[Y∣X]−E[Y]∥22​]​
This measure takes values in [0,1], vanishing if and only if Y is conditionally mean independent of X, and attaining $1$ if Y is a.s. a function of X. For scalar Y∈Rp0, Y∈Rp1 is the classical Sobol' index Y∈Rp2. For vector-valued responses, the natural extension is the trace ratio form, connecting the measure to multivariate GSA [Gamboa et al.].
The relationship between conditional mean independence and Sobol' indices provides direct bridges between model diagnostics, feature relevance, and the sensitivity of model outputs to input variables.
Nearest Neighbor Graph Estimators: Construction, Consistency, and Rates
The NN graph-based estimator exploits the proximity structure of the covariate space. For i.i.d. samples Y∈Rp3, the Y∈Rp4-NN graph on Y∈Rp5 determines sets of local neighborhoods. The numerator of the estimator computes an average inner product between a point's response and those of its NN neighbors, with bias corrections corresponding to the empirical mean.
The proposed estimator satisfies key properties:
- Computational Efficiency: For fixed Y∈Rp6, it is computable in Y∈Rp7 time, making it scalable to high dimensions and large Y∈Rp8.
- Consistency and Rate: Strong uniform consistency is established under mild moment and smoothness conditions on the regression function. The rate is Y∈Rp9, i.e., nearly parametric for low-dimensional covariates, with bias dominating for X∈Rd0 due to curse-of-dimensionality phenomena inherent in all nonparametric dependence estimation.
Asymptotic Testing for Conditional Mean Independence
A major contribution is the derivation of an asymptotically normal, variance-stabilized test statistic for conditional mean independence, requiring no resampling or sample splitting. Importantly, the asymptotic distribution is Gaussian under the null hypothesis for arbitrary dimensions, enabling analytical threshold selection.
Theoretical results ensure universal consistency: under the null, the type I error matches the nominal level asymptotically; for any fixed alternative, the power converges to one. The variance estimator for normalization is fully data-driven, relying only on observed residuals and graph structure.

Figure 1: Empirical Type I error, power, and computational time for conditional mean independence testing on univariate settings, showing tight type I control and advantageous power/computation balance for the NN method.
Empirical evaluation (see Figure 1) shows:
- The NN graph test substantially outperforms kernel-based, martingale-difference, and split-sample machine learning approaches in computational efficiency.
- The type I error is precisely controlled in null models satisfying conditional mean independence but violating unconditional independence, where classical independence tests such as dCov fail.
- Superior or competitive power across a suite of complex, nonlinear alternatives, including additive, interaction, and heteroskedastic models.
Model-Free Variable Screening Algorithm
The NN estimator permits a forward model-free variable screening procedure, selecting variables in order of their marginal increase to the estimated mean dependence index. At each step, the variable providing maximal increase to the estimated NCMD is added, guaranteeing under mild conditions that the resulting subset is sufficiently informative with exponentially small probability of error as X∈Rd1 grows.
In high-dimensional simulated variable selection benchmarks, the procedure is competitive or superior to state-of-the-art model-free screening algorithms, such as KFOCI, BcorSIS, and MDCSIS.
Higher-Order Sobol' Indices and Interaction Effects
The framework extends naturally to estimation of higher-order Sobol' indices, allowing quantification of pure interaction effects among groups of variables beyond individual (main) effects. Specifically, second-order indices are efficiently estimated using NN graphs constructed on the joint subspace of the interacting variables, with correction for marginal effects.

Figure 2: Estimated and true Sobol' indices (main and interaction effects) as a function of model parameter X∈Rd2, demonstrating tight alignment and convergence as sample size increases.
This enables detailed variance decomposition even in multivariate models, with applications encompassing interaction detection and GSA in complex simulation models.
Practical and Theoretical Implications
The proposed NN graph methodology bridges nonparametric regression diagnostics, feature selection, and global sensitivity analysis under a unified computational and statistical paradigm:
- Practical: It enables rapid, distribution-free variable importance and independence assessment, scalable to high dimensions, with immediate applicability to black-box surrogates and simulation experiments prevalent in scientific computation and uncertainty quantification.
- Theoretical: The results clarify the connections between classical variance-based sensitivity measures and modern dependence testing, while demonstrating the power of geometric, proximity-based statistics for fundamental inference tasks.
- Broader AI Impact: This methodology has implications for explainable AI and interpretable machine learning, especially in black-box contexts (e.g., deep or ensemble models) where parametric model assumptions are untenable. The ability to efficiently assess variable relevance and interaction without resampling or splitting is critical in high-throughput or large-scale AI pipelines.
Conclusion
This work establishes nearest neighbor graph-based estimation as a powerful and unified tool for nonparametric conditional mean inference, variable selection, and sensitivity analysis. The strong theoretical guarantees, computational scalability, and empirical effectiveness position the methods as valuable for both statistical methodology and applied AI applications requiring model-free assessment of covariate influence and sensitivity structure. Future developments may address further reduction of bias in very high dimensions, direct incorporation with adversarial robustness analysis, and applications to time series or dependent data settings.