Papers
Topics
Authors
Recent
Search
2000 character limit reached

Non-asymptotic two-sample kernel testing with the spectrally truncated normalized MMD

Published 8 Apr 2026 in math.ST and stat.ME | (2604.07153v1)

Abstract: Kernel methods provide a flexible and powerful framework for nonparametric statistical testing by embedding probability distributions into a reproducing kernel Hilbert space (RKHS). In this work, we study the kernel two-sample testing problem and focus on a normalized version of the Maximum Mean Discrepancy (MMD) as a test statistic, which scales the discrepancy by the within-group covariance operator to account for data variability. This normalization has been shown to improve test power in both theoretical and empirical settings. Because this normalization requires regularization, we study the non-asymptotic properties of the spectrally truncated normalized MMD (st-nMMD) and derive an exponential upper bound under the null hypothesis. Thanks to this result we propose a sharp and explicit upper bound for the corresponding non-asymptotic quantile, along with a data-adaptive estimator. We further propose an algorithm to tune the hyperparameters involved in the quantile estimation, including the truncation level, without requiring data splitting. We demonstrate the performance of the st-nMMD through numerical experiments under both the null and alternative hypotheses.

Summary

  • The paper presents a spectrally truncated normalized MMD test that uses within-group covariance operators to achieve sharp non-asymptotic error control.
  • It introduces a fully data-adaptive calibration procedure that selects the spectral truncation level without data splitting, balancing power and conservatism.
  • Empirical results on simulated data and MNIST validate the test’s competitive performance and asymptotic consistency even in high-dimensional, small-sample contexts.

Non-Asymptotic Kernel Two-Sample Testing via Spectrally Truncated Normalized MMD

Introduction

The kernel two-sample testing problem has seen considerable advances via the Maximum Mean Discrepancy (MMD) statistic, which interprets the equality of distributions problem in terms of the norm of mean differences in a Reproducing Kernel Hilbert Space (RKHS). However, classical MMD-based tests relying on asymptotic approximations often encounter calibration challenges in finite-sample, high-dimensional contexts. This paper addresses these shortcomings by proposing and rigorously analyzing the spectrally truncated normalized MMD (st-nMMD) test, which normalizes the mean embedding difference by the within-group covariance operator and regularizes via spectral truncation, enabling theoretically sharp, non-asymptotic error control.

Statistical Formulation and Rationale

The st-nMMD statistic is defined as:

D^T2=nXnYnX+nY∑t=1T⟨f^t,μ^X−μ^Y⟩H2λ^t\widehat{D}^2_T = \frac{n_X n_Y}{n_X + n_Y} \sum_{t=1}^T \frac{\langle \widehat{f}_t, \widehat{\mu}_X - \widehat{\mu}_Y\rangle_H^2}{\widehat{\lambda}_t}

where Σ^\widehat{\Sigma} is the empirical within-group covariance operator, λ^t\widehat{\lambda}_t and f^t\widehat{f}_t are its spectral components, and TT is the spectral truncation level. By incorporating spectral directions associated with significant variance and appropriately regularizing via truncation, this statistic achieves both sensitivity to distributional differences and robust finite-sample behavior.

Non-Asymptotic Quantile Calibration

A central contribution is the derivation of finite-sample exponential deviation inequalities for the st-nMMD statistic under the null hypothesis, leveraging concentration inequalities for self-normalized processes in Hilbert spaces. The result provides explicit, computable upper bounds for the quantiles of D^T2\widehat{D}^2_T that avoid data splitting and enable sharp type-I error control even in high dimensions and for small to moderate sample sizes. The derivations crucially depend on lower bounds for spectral gaps and eigenvalues, ensuring stable estimation of discriminant directions.

Asymptotic Consistency and Optimality

In the asymptotic regime (n→∞n\rightarrow\infty), the st-nMMD statistic converges in distribution to a χ2(T)\chi^2(T) law, aligning with the classical Hotelling T2T^2 kernelized test and guaranteeing consistency. The non-asymptotic quantile bounds recover, up to multiplicative constants, the optimal asymptotic quantiles, establishing minimax optimality against smooth alternatives under regularity assumptions. Explicit scaling between eigenvalue decay, spectral gap size, and sample size is analyzed in detail—yielding clear practical guidance for balancing statistical power and calibration.

Data-Adaptive Calibration and Practical Algorithm

A fully data-driven algorithm is introduced for the selection of spectral truncation level TT and calibration of quantile hyperparameters. The method operates without data splitting and ensures computational tractability by simplifying the quantile formula and tuning constants directly on the observed data. This yields a test procedure that adapts not only to sample size and dimensionality but also to the empirical covariance structure, thus achieving robust calibration and competitive power across a wide range of simulation and real-world scenarios.

Empirical Evaluation

Comprehensive numerical experiments are conducted on simulated isotropic distributions (Gaussian, uniform, Cauchy, von Mises–Fisher) across a range of sample sizes and dimensions, as well as on the MNIST dataset, comparing the proposed procedure to classical Σ^\widehat{\Sigma}0-based thresholds. Empirical levels for st-nMMD consistently align with nominal type-I error rates, even for small Σ^\widehat{\Sigma}1 and low Σ^\widehat{\Sigma}2, while asymptotic calibration (based on Σ^\widehat{\Sigma}3) tends to underperform, especially in finite-sample regimes. Figure 1

Figure 1: Average empirical level of st-nMMD for varying truncation Σ^\widehat{\Sigma}4, showing strong calibration of the non-asymptotic quantile relative to the Σ^\widehat{\Sigma}5 approach across multiple distributions and dimensions.

Figure 2

Figure 2: Empirical level at Σ^\widehat{\Sigma}6, confirming calibration robustness in more stringent testing settings.

Power analysis reveals that data-adaptive quantile calibration does not sacrifice sensitivity under alternatives—even in scenarios where conservative level control might intuitively suggest loss of power. Adaptive tuning of hyperparameters, including Σ^\widehat{\Sigma}7, enables balancing conservatism and power. Figure 3

Figure 3: Empirical level for MNIST (Σ^\widehat{\Sigma}8); non-asymptotic quantile maintains tight control.

Figure 4

Figure 4: Empirical power on MNIST, illustrating competitive performance of the st-nMMD test with data-adaptive quantiles across varying alternatives and truncation levels.

Automatic selection of the truncation parameter Σ^\widehat{\Sigma}9 is analyzed via frequency histograms under both null and alternative scenarios, demonstrating that selection remains reliably low unless justified by observed structure, thereby preventing overfitting. Figure 5

Figure 5: Frequency distribution of selected λ^t\widehat{\lambda}_t0 in simulated null samples.

Figure 6

Figure 6: Selected λ^t\widehat{\lambda}_t1 frequencies under MNIST null.

Figure 7

Figure 7: Selected λ^t\widehat{\lambda}_t2 frequencies under MNIST alternative; modest growth of λ^t\widehat{\lambda}_t3 with increased sample size and difference complexity.

Tests with automatically selected λ^t\widehat{\lambda}_t4 consistently exhibit calibrated error rates; empirical power remains high even with conservative quantile adjustments. Figure 8

Figure 8: Empirical level for algorithmically selected λ^t\widehat{\lambda}_t5 across distributions and dimensions.

Figure 9

Figure 9: Empirical level for selected λ^t\widehat{\lambda}_t6 on MNIST null.

Figure 10

Figure 10: Empirical power for selected λ^t\widehat{\lambda}_t7 on MNIST alternatives; power approaches unity for large λ^t\widehat{\lambda}_t8.

Theoretical and Practical Implications

The st-nMMD test establishes a new standard for kernel two-sample testing in non-asymptotic settings, particularly in high-dimensional and low-sample scenarios. By making the quantile calibration data-adaptive—and avoiding data splitting—practitioners gain a tractable, theoretically rigorous method that leverages spectral information optimally. The dependence on empirical covariance structure enables richer interpretations and paves the way for principled integration with nonlinear data representation and visualization tasks. Extensions to settings with unbalanced sample sizes and multiplicities in covariance spectra are immediate, albeit with increased analytical complexity.

Conclusion

The spectrally truncated normalized MMD test, with sharp non-asymptotic calibration and fully data-adaptive implementation, resolves key limitations of classical kernel-based two-sample testing. Across theoretical and empirical axes, the proposed approach offers robust level control, competitive power, and flexible regularization, positioning it as a practical and principled procedure for high-dimensional, nonparametric inferential tasks. Future methodological developments will deepen the minimax power analysis, relax balanced-sample size assumptions, and further integrate representational directions with interpretable testing frameworks.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.