Papers
Topics
Authors
Recent
Search
2000 character limit reached

Adversarial Robustness of NTK Neural Networks

Published 28 Apr 2026 in stat.ML and cs.LG | (2604.25965v1)

Abstract: Deep learning models are widely deployed in safety-critical domains, but remain vulnerable to adversarial attacks. In this paper, we study the adversarial robustness of NTK neural networks in the context of nonparametric regression. We establish minimax optimal rates for adversarial regression in Sobolev spaces and then show that NTK neural networks, trained via gradient flow with early stopping, can achieve this optimal rate. However, in the overfitting regime, we prove that the minimum norm interpolant is vulnerable to adversarial perturbations.

Authors (1)

Summary

  • The paper presents a theoretical framework that derives the minimax optimal rate for adversarial risk using NTK analysis and Sobolev space regularity.
  • It demonstrates that early stopping in gradient flow training is crucial to achieving adversarial optimality and preventing overfit-induced instabilities.
  • The study shows that overfitting, contrary to benign overfitting in standard regression, induces significant adversarial vulnerability in the NTK regime.

Adversarial Robustness of Wide Neural Networks in the NTK Regime

Introduction

The study examines the adversarial robustness properties of neural networks in the Neural Tangent Kernel (NTK) regime. It rigorously characterizes statistical limits for adversarially robust nonparametric regression, provides an adversarial risk analysis of infinitely wide fully connected ReLU networks, and interrogates the effects of overfitting with respect to adversarial risk. The framework offers a unified treatment of minimax optimality under adversarial perturbations, connects network dynamics with Sobolev regularity, and delivers negative results regarding the "benign overfitting" phenomenon in the adversarial context.

Problem Setup and Theoretical Foundations

The paper operates within a standard nonparametric regression model:

Yi=f(Xi)+ξi,Xi[0,1]d,Y_i = f^*(X_i) + \xi_i, \quad X_i \in [0,1]^d,

where ff^* lies in the Sobolev space Hs([0,1]d)H^s([0,1]^d), and ξi\xi_i is sub-Gaussian noise. The adversarial risk is defined as the expected supremal squared error over rr-ball perturbations to each input, quantifying the worst-case vulnerability to norm-bounded attacks.

The principal technical contribution in this setup is the derivation of the minimax optimal rate for regression under adversarial risk:

inff^supfHs(L)RA(f^,f)r2min(1,s)+n2s2s+d,\inf_{\hat{f}} \sup_{f^* \in \mathcal{H}^s(L)} R_A(\hat{f}, f^*) \asymp r^{2\min(1,s)} + n^{-\frac{2s}{2s+d}},

where r2min(1,s)r^{2\min(1,s)} captures adversarial instability and n2s2s+dn^{-\frac{2s}{2s+d}} is the minimax estimation error without adversarial considerations. This result delineates the fundamental statistical limit, separating approximation imposed by adversarial radius from sample complexity.

The NTK Regime and Achievability

Using mirrored, fully connected ReLU architectures in the infinite-width (NTK) regime, the authors connect gradient flow training dynamics to predictions in the corresponding RKHS. The NTK on [0,1]d[0,1]^d is norm-equivalent to the Sobolev space Hd+12([0,1]d)H^{\frac{d+1}{2}}([0,1]^d), establishing the interface between network architecture, expressivity, and regularity.

The main achievability result demonstrates that, with early stopping at ff^*0, wide ReLU networks attain:

ff^*1

for ff^*2 with ff^*3. Early stopping regularizes the gradient flow to cap the Sobolev norm of the estimator, preventing overfit-induced oscillations, and ensures adversarial optimality up to constants. This result aligns the practical procedure (early-stopped gradient descent) with the theoretical minimax rate, a notable analytical closure in adversarial statistics and infinite-width neural networks.

Failures of Benign Overfitting for Adversarial Robustness

A major finding is a sharp negative result on overparametrized regime behavior under adversarial risk. In contrast to "benign overfitting" phenomena in standard (non-adversarial) regression—where fits to noisy data generalize well in expectation—the paper proves that fitting to stochastic noise is fundamentally incompatible with adversarial robustness. For the minimum norm interpolant (gradient flow as ff^*4),

ff^*5

where ff^*6 is the sequence of perturbation radii. This lower bound, growing logarithmically in ff^*7, demonstrates unavoidable degradation in adversarial risk as the estimator develops high-frequency, non-robust features to interpolate noise. The result strictly demarcates the limits of overfitting tolerance: perfect data-fitting induces adversarial brittleness even as standard risk remains small.

Empirical Evidence

Empirical validation using both synthetic and real data (including the Diabetes dataset) supports the theoretical insights. During the descent phase of gradient flow, adversarial risk aligns with theory, declining with sample complexity and controlled adversarial radius. As training proceeds to interpolation (ascent phase), adversarial risk sharply increases, visuals confirming the theoretical claim that accelerated adversarial vulnerability emerges with overfitting.

Implications and Future Directions

This study establishes that, at least in the NTK regime, the achievable adversarial risk is fundamentally characterized by the interplay of Sobolev regularity, adversarial perturbation scale, and sample size. The convergence between the NTK and minimax rates confirms that, under appropriate regularization (early stopping), wide neural networks are statistically optimal for adversarially robust learning.

On the negative side, the sharp separation from benign overfitting underscores a need to rethink model selection and regularization in safety-critical applications, as interpolation need not coincide with robustness.

Natural follow-ups include:

  • Generalization to ff^*8 adversarial losses, beyond ff^*9,
  • Transcending the NTK limit to incorporate finite-width or feature learning effects,
  • Characterizing transitions and limits as Hs([0,1]d)H^s([0,1]^d)0 with growing Hs([0,1]d)H^s([0,1]^d)1.

These directions open space for future theoretical and practical advances in robust AI.

Conclusion

The paper delivers a comprehensive statistical analysis of adversarial robustness in wide neural networks, precisely connects the NTK regime to minimax adversarial risk via early stopping, and proves the incompatibility of benign overfitting with adversarially robust learning. These results contribute important precision to the theoretical landscape, guiding both practice (regarding regularization) and future research in adversarially robust deep learning.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 1 tweet with 2 likes about this paper.