Papers
Topics
Authors
Recent
Search
2000 character limit reached

Natural Wasserstein Metric for Robust Attribution

Updated 11 December 2025
  • Natural Wasserstein Metric is a data-dependent ground metric that integrates learned feature covariance to stabilize attribution estimates in machine learning models.
  • It replaces the Euclidean measure with a Mahalanobis-type metric, significantly reducing spectral amplification and enabling non-vacuous robust certification in deep networks.
  • The approach combines robust influence function analysis, sensitivity kernel evaluation, and Lipschitz constant estimation to deliver efficient, closed-form influence intervals in both convex and deep models.

The Natural Wasserstein metric is a data-dependent ground metric that quantifies perturbations in the geometry induced by a model’s own feature covariance. It is designed to address the problem of spectral amplification in distributionally robust data attribution, stabilizing attribution estimates and enabling non-vacuous certification for influence-based data attribution methods in both convex models and deep neural networks (Li et al., 9 Dec 2025).

1. Influence Functions and Distributional Robustness

Classical influence functions quantify the infinitesimal effect of upweighting a single training example ziz_i on the prediction or loss for a test input ztestz_{\rm test}. In the context of empirical risk minimization,

I(zi,ztest)=−gtest⊤H−1gi\mathcal I(z_i, z_{\rm test}) = -g_{\rm test}^\top H^{-1} g_i

where gi=∇θℓ(θ^,zi)g_i = \nabla_\theta \ell(\hat\theta, z_i), gtest=∇θℓ(θ^,ztest)g_{\rm test} = \nabla_\theta \ell(\hat\theta, z_{\rm test}), and H=1n∑j=1n∇θ2ℓ(θ^,zj)H = \frac{1}{n} \sum_{j=1}^n \nabla^2_\theta \ell(\hat\theta, z_j) with ℓ(θ,z)\ell(\theta, z) a C3C^3 loss and H≻0H \succ 0.

However, standard attribution scores are highly sensitive to distributional perturbations. The Wasserstein-Robust Influence Function (W-RIF) formalizes this by taking the supremum of the influence functional over all distributions QQ within a ztestz_{\rm test}0-Wasserstein ball of radius ztestz_{\rm test}1 centered at the empirical distribution ztestz_{\rm test}2: ztestz_{\rm test}3 This formulation rigorously certifies attribution stability against worst-case distributional drifts (Li et al., 9 Dec 2025).

2. Sensitivity Kernel and Robust Influence Certificates

W-RIF can be analyzed through the first-order expansion of the influence functional in powers of ztestz_{\rm test}4: ztestz_{\rm test}5 where the sensitivity kernel ztestz_{\rm test}6 is

ztestz_{\rm test}7

with ztestz_{\rm test}8, ztestz_{\rm test}9, I(zi,ztest)=−gtest⊤H−1gi\mathcal I(z_i, z_{\rm test}) = -g_{\rm test}^\top H^{-1} g_i0, I(zi,ztest)=−gtest⊤H−1gi\mathcal I(z_i, z_{\rm test}) = -g_{\rm test}^\top H^{-1} g_i1, I(zi,ztest)=−gtest⊤H−1gi\mathcal I(z_i, z_{\rm test}) = -g_{\rm test}^\top H^{-1} g_i2.

By applying Kantorovich–Rubinstein duality, the robustness certificate is governed by the Lipschitz constant I(zi,ztest)=−gtest⊤H−1gi\mathcal I(z_i, z_{\rm test}) = -g_{\rm test}^\top H^{-1} g_i3 of I(zi,ztest)=−gtest⊤H−1gi\mathcal I(z_i, z_{\rm test}) = -g_{\rm test}^\top H^{-1} g_i4: I(zi,ztest)=−gtest⊤H−1gi\mathcal I(z_i, z_{\rm test}) = -g_{\rm test}^\top H^{-1} g_i5 The closed-form certificate for small I(zi,ztest)=−gtest⊤H−1gi\mathcal I(z_i, z_{\rm test}) = -g_{\rm test}^\top H^{-1} g_i6 is: I(zi,ztest)=−gtest⊤H−1gi\mathcal I(z_i, z_{\rm test}) = -g_{\rm test}^\top H^{-1} g_i7 yielding efficient, closed-form robust influence intervals in convex settings (Li et al., 9 Dec 2025).

3. Natural Wasserstein Metric and Spectral Amplification

In deep neural networks, naïvely applying a Euclidean Wasserstein ball in high-dimensional feature space yields vacuous certifications due to spectral amplification: the ill-conditioning of learned feature covariance matrices inflates Lipschitz bounds by factors exceeding I(zi,ztest)=−gtest⊤H−1gi\mathcal I(z_i, z_{\rm test}) = -g_{\rm test}^\top H^{-1} g_i8. This is observed empirically as 0% certified robustness for methods such as TRAK when using Euclidean geometry (Li et al., 9 Dec 2025).

The Natural Wasserstein metric replaces the ground metric I(zi,ztest)=−gtest⊤H−1gi\mathcal I(z_i, z_{\rm test}) = -g_{\rm test}^\top H^{-1} g_i9 with the Mahalanobis-type metric gi=∇θℓ(θ^,zi)g_i = \nabla_\theta \ell(\hat\theta, z_i)0, where gi=∇θℓ(θ^,zi)g_i = \nabla_\theta \ell(\hat\theta, z_i)1 denotes the feature covariance. This alignment with the geometry of the learned representation eliminates spectral amplification and ensures that the distributional perturbation set matches the curvature of the sensitivity kernel.

Empirical results demonstrate that replacing the Euclidean metric with the Natural metric reduces worst-case sensitivity by gi=∇θℓ(θ^,zi)g_i = \nabla_\theta \ell(\hat\theta, z_i)2, yielding non-vacuous, certifiable robust attribution for deep networks. For example, using ResNet-18 on CIFAR-10, Natural W-TRAK certifies 68.7% of ranking pairs, compared to 0% for the Euclidean baseline (Li et al., 9 Dec 2025).

4. Algorithmic Procedure and Complexity

The W-RIF pipeline in the convex setting involves:

  1. Solving the ERM to obtain gi=∇θℓ(θ^,zi)g_i = \nabla_\theta \ell(\hat\theta, z_i)3.
  2. Forming and inverting the Hessian gi=∇θℓ(θ^,zi)g_i = \nabla_\theta \ell(\hat\theta, z_i)4.
  3. Computing gradients and nominal influences for each training point.
  4. Evaluating the sensitivity kernel gi=∇θℓ(θ^,zi)g_i = \nabla_\theta \ell(\hat\theta, z_i)5 for each gi=∇θℓ(θ^,zi)g_i = \nabla_\theta \ell(\hat\theta, z_i)6.
  5. Estimating the Lipschitz constant gi=∇θℓ(θ^,zi)g_i = \nabla_\theta \ell(\hat\theta, z_i)7 by maximizing gi=∇θℓ(θ^,zi)g_i = \nabla_\theta \ell(\hat\theta, z_i)8.
  6. Forming robust intervals gi=∇θℓ(θ^,zi)g_i = \nabla_\theta \ell(\hat\theta, z_i)9.

Computational costs are gtest=∇θℓ(θ^,ztest)g_{\rm test} = \nabla_\theta \ell(\hat\theta, z_{\rm test})0 for gtest=∇θℓ(θ^,ztest)g_{\rm test} = \nabla_\theta \ell(\hat\theta, z_{\rm test})1 and gtest=∇θℓ(θ^,ztest)g_{\rm test} = \nabla_\theta \ell(\hat\theta, z_{\rm test})2, gtest=∇θℓ(θ^,ztest)g_{\rm test} = \nabla_\theta \ell(\hat\theta, z_{\rm test})3 for gtest=∇θℓ(θ^,ztest)g_{\rm test} = \nabla_\theta \ell(\hat\theta, z_{\rm test})4, and gtest=∇θℓ(θ^,ztest)g_{\rm test} = \nabla_\theta \ell(\hat\theta, z_{\rm test})5 for pairwise Lipschitz calculation (with gtest=∇θℓ(θ^,ztest)g_{\rm test} = \nabla_\theta \ell(\hat\theta, z_{\rm test})6 the cost of the ground metric). In high dimensions, gtest=∇θℓ(θ^,ztest)g_{\rm test} = \nabla_\theta \ell(\hat\theta, z_{\rm test})7 may be estimated from random sample pairs or local neighborhoods (Li et al., 9 Dec 2025).

5. Theoretical Guarantees and Duality

Coverage guarantees for robust influence intervals are established via concentration inequalities for empirical Wasserstein distances. With appropriate adjustment of the radius by the empirical deviation gtest=∇θℓ(θ^,ztest)g_{\rm test} = \nabla_\theta \ell(\hat\theta, z_{\rm test})8, the interval

gtest=∇θℓ(θ^,ztest)g_{\rm test} = \nabla_\theta \ell(\hat\theta, z_{\rm test})9

contains the true population-level robust influence with high probability. The derivation leverages distributionally robust optimization duality and Kantorovich–Rubinstein duality, ensuring that the certificate reflects the worst-case influence achievable under feasible Wasserstein perturbations (Li et al., 9 Dec 2025).

6. Extensions to Deep Networks and Attribution Stability

For deep neural networks, the convex W-RIF approach does not directly transfer due to nonconvexity and initialization sensitivity ("basin-hopping" under distribution shifts). The adopted strategy is to linearize the influence functional in feature space for the fixed (pretrained or converged) network, then apply the Wasserstein DRO analysis using the Natural metric. This exactly compensates for feature-space ill-conditioning.

The Self-Influence term is shown to equal the Lipschitz constant governing attribution stability, which provides a theoretical foundation for leverage-based anomaly detection. Empirically, Self-Influence yields strong results for label noise detection, achieving 0.970 AUROC and identifying 94.1% of corrupted labels within the top 20% of training data (Li et al., 9 Dec 2025).

7. Summary and Significance

The Natural Wasserstein metric constitutes a principled, geometrically-adaptive ground metric for robust data attribution. By integrating feature covariance structure, it removes the instability inherent to Euclidean robustness analysis in high-dimensional (especially deep) models and enables the first non-vacuous certified bounds for neural network attribution. The approach generalizes from convex models to deep networks, systematically aligning the perturbation geometry with the local sensitivity landscape, thereby producing efficient, closed-form, and certifiably tight influence intervals fundamental for robust, interpretable machine learning (Li et al., 9 Dec 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Natural Wasserstein Metric.