Non-asymptotic two-sample kernel testing with the spectrally truncated normalized MMD
Published 8 Apr 2026 in math.ST and stat.ME | (2604.07153v1)
Abstract: Kernel methods provide a flexible and powerful framework for nonparametric statistical testing by embedding probability distributions into a reproducing kernel Hilbert space (RKHS). In this work, we study the kernel two-sample testing problem and focus on a normalized version of the Maximum Mean Discrepancy (MMD) as a test statistic, which scales the discrepancy by the within-group covariance operator to account for data variability. This normalization has been shown to improve test power in both theoretical and empirical settings. Because this normalization requires regularization, we study the non-asymptotic properties of the spectrally truncated normalized MMD (st-nMMD) and derive an exponential upper bound under the null hypothesis. Thanks to this result we propose a sharp and explicit upper bound for the corresponding non-asymptotic quantile, along with a data-adaptive estimator. We further propose an algorithm to tune the hyperparameters involved in the quantile estimation, including the truncation level, without requiring data splitting. We demonstrate the performance of the st-nMMD through numerical experiments under both the null and alternative hypotheses.
The paper presents a spectrally truncated normalized MMD test that uses within-group covariance operators to achieve sharp non-asymptotic error control.
It introduces a fully data-adaptive calibration procedure that selects the spectral truncation level without data splitting, balancing power and conservatism.
Empirical results on simulated data and MNIST validate the test’s competitive performance and asymptotic consistency even in high-dimensional, small-sample contexts.
Non-Asymptotic Kernel Two-Sample Testing via Spectrally Truncated Normalized MMD
Introduction
The kernel two-sample testing problem has seen considerable advances via the Maximum Mean Discrepancy (MMD) statistic, which interprets the equality of distributions problem in terms of the norm of mean differences in a Reproducing Kernel Hilbert Space (RKHS). However, classical MMD-based tests relying on asymptotic approximations often encounter calibration challenges in finite-sample, high-dimensional contexts. This paper addresses these shortcomings by proposing and rigorously analyzing the spectrally truncated normalized MMD (st-nMMD) test, which normalizes the mean embedding difference by the within-group covariance operator and regularizes via spectral truncation, enabling theoretically sharp, non-asymptotic error control.
where Σ is the empirical within-group covariance operator, λt​ and f​t​ are its spectral components, and T is the spectral truncation level. By incorporating spectral directions associated with significant variance and appropriately regularizing via truncation, this statistic achieves both sensitivity to distributional differences and robust finite-sample behavior.
Non-Asymptotic Quantile Calibration
A central contribution is the derivation of finite-sample exponential deviation inequalities for the st-nMMD statistic under the null hypothesis, leveraging concentration inequalities for self-normalized processes in Hilbert spaces. The result provides explicit, computable upper bounds for the quantiles of DT2​ that avoid data splitting and enable sharp type-I error control even in high dimensions and for small to moderate sample sizes. The derivations crucially depend on lower bounds for spectral gaps and eigenvalues, ensuring stable estimation of discriminant directions.
Asymptotic Consistency and Optimality
In the asymptotic regime (n→∞), the st-nMMD statistic converges in distribution to a χ2(T) law, aligning with the classical Hotelling T2 kernelized test and guaranteeing consistency. The non-asymptotic quantile bounds recover, up to multiplicative constants, the optimal asymptotic quantiles, establishing minimax optimality against smooth alternatives under regularity assumptions. Explicit scaling between eigenvalue decay, spectral gap size, and sample size is analyzed in detail—yielding clear practical guidance for balancing statistical power and calibration.
Data-Adaptive Calibration and Practical Algorithm
A fully data-driven algorithm is introduced for the selection of spectral truncation level T and calibration of quantile hyperparameters. The method operates without data splitting and ensures computational tractability by simplifying the quantile formula and tuning constants directly on the observed data. This yields a test procedure that adapts not only to sample size and dimensionality but also to the empirical covariance structure, thus achieving robust calibration and competitive power across a wide range of simulation and real-world scenarios.
Empirical Evaluation
Comprehensive numerical experiments are conducted on simulated isotropic distributions (Gaussian, uniform, Cauchy, von Mises–Fisher) across a range of sample sizes and dimensions, as well as on the MNIST dataset, comparing the proposed procedure to classical Σ0-based thresholds. Empirical levels for st-nMMD consistently align with nominal type-I error rates, even for small Σ1 and low Σ2, while asymptotic calibration (based on Σ3) tends to underperform, especially in finite-sample regimes.
Figure 1: Average empirical level of st-nMMD for varying truncation Σ4, showing strong calibration of the non-asymptotic quantile relative to the Σ5 approach across multiple distributions and dimensions.
Figure 2: Empirical level at Σ6, confirming calibration robustness in more stringent testing settings.
Power analysis reveals that data-adaptive quantile calibration does not sacrifice sensitivity under alternatives—even in scenarios where conservative level control might intuitively suggest loss of power. Adaptive tuning of hyperparameters, including Σ7, enables balancing conservatism and power.
Figure 4: Empirical power on MNIST, illustrating competitive performance of the st-nMMD test with data-adaptive quantiles across varying alternatives and truncation levels.
Automatic selection of the truncation parameter Σ9 is analyzed via frequency histograms under both null and alternative scenarios, demonstrating that selection remains reliably low unless justified by observed structure, thereby preventing overfitting.
Figure 5: Frequency distribution of selected λt​0 in simulated null samples.
Figure 6: Selected λt​1 frequencies under MNIST null.
Figure 7: Selected λt​2 frequencies under MNIST alternative; modest growth of λt​3 with increased sample size and difference complexity.
Tests with automatically selected λt​4 consistently exhibit calibrated error rates; empirical power remains high even with conservative quantile adjustments.
Figure 8: Empirical level for algorithmically selected λt​5 across distributions and dimensions.
Figure 9: Empirical level for selected λt​6 on MNIST null.
Figure 10: Empirical power for selected λt​7 on MNIST alternatives; power approaches unity for large λt​8.
Theoretical and Practical Implications
The st-nMMD test establishes a new standard for kernel two-sample testing in non-asymptotic settings, particularly in high-dimensional and low-sample scenarios. By making the quantile calibration data-adaptive—and avoiding data splitting—practitioners gain a tractable, theoretically rigorous method that leverages spectral information optimally. The dependence on empirical covariance structure enables richer interpretations and paves the way for principled integration with nonlinear data representation and visualization tasks. Extensions to settings with unbalanced sample sizes and multiplicities in covariance spectra are immediate, albeit with increased analytical complexity.
Conclusion
The spectrally truncated normalized MMD test, with sharp non-asymptotic calibration and fully data-adaptive implementation, resolves key limitations of classical kernel-based two-sample testing. Across theoretical and empirical axes, the proposed approach offers robust level control, competitive power, and flexible regularization, positioning it as a practical and principled procedure for high-dimensional, nonparametric inferential tasks. Future methodological developments will deepen the minimax power analysis, relax balanced-sample size assumptions, and further integrate representational directions with interpretable testing frameworks.
“Emergent Mind helps me see which AI papers have caught fire online.”
Philip
Creator, AI Explained on YouTube
Sign up for free to explore the frontiers of research
Discover trending papers, chat with arXiv, and track the latest research shaping the future of science and technology.Discover trending papers, chat with arXiv, and more.