- The paper introduces DAVE, a method that decomposes ViT gradients to isolate effective input–output transformations while filtering out architectural noise.
- It applies spatial perturbations and low-pass filtering to generate object-centric, stable attribution maps with consistent class-level features.
- DAVE outperforms previous methods, achieving substantial gains in localization, faithfulness, and robustness on benchmarks like ImageNet-1k.
DAVE: Distribution-aware Attribution via ViT Gradient Decomposition
The proliferation of Vision Transformers (ViTs) in computer vision necessitates explainability methods that can generate fine-grained, robust, and class-consistent attribution maps. The inherent architecture of ViTs—patch embedding, multi-head attention, nonlinearity, and normalization—introduces structured artifacts and instability in pixel-level attributions. Previous attribution methods, both generic (e.g., Input×Gradient, Integrated Gradients, SmoothGrad) and ViT-specific (e.g., AttnLRP, LeGrad, Chefer-LRP) have struggled to suppress such artifacts, frequently resorting to coarse, patch-level explanations or producing maps highly sensitive to input perturbations.
The DAVE Framework: Mathematical Foundations and Pipeline
DAVE ("Distribution-aware Attribution via ViT Gradient Decomposition") proposes a mathematically principled, architecture-aware approach for ViT model attribution. The method starts by explicitly decomposing the input gradient of a ViT into two components: the effective input–output transformation (capturing direct dependence on the input) and the operator-variation term (capturing input-sensitive architectural noise and instabilities). By isolating the effective transformation and discarding the operator-variation term, DAVE extracts the baseline attribution signal with reduced instability.
To further suppress architecture-induced artifacts (grid patterns, attention patch effects), DAVE leverages a Reynolds-inspired operator that aggregates attribution maps over small spatial transformations (rotations, translations) and Gaussian perturbations, retaining only locally equivariant features. Finally, DAVE applies a low-pass filter by averaging attribution maps over input noise, yielding high stability and spatial coherence.

Figure 1: Overview of the DAVE attribution pipeline, including spatial transformation/perturbation sampling, effective input–output transformation extraction, operator variation filtering, aggregation, and final attribution map computation.

Figure 2: Progressive construction of DAVE attribution, showing input, inputĂ—gradient, inputĂ—effective transformation, inputĂ—equivariant transformation, and the final DAVE attribution map with suppressed artifacts and improved interpretability.
Attribution Consistency and Faithfulness
DAVE enforces attribution consistency under small augmentations (e.g., rotation, shifts), ensuring that the same semantically relevant features are highlighted, even as the input changes. Empirically, prior methods such as AttnLRP and LeGrad yield inconsistent attributions for augmented inputs (marked by white regions), whereas DAVE provides stable, consistent feature localization.

Figure 3: DAVE attribution consistency under small spatial augmentations; DAVE highlights stable features while prior methods yield inconsistent explanations.
Faithfulness is validated via pixel deletion experiments: as pixels ranked least important by the attribution map are progressively removed, DAVE maintains target-class probabilities longer than competing methods, producing the flattest deletion curves and highest AUC, confirming that DAVE explanations reliably correspond to actual model behavior.


Figure 4: Pixel deletion faithfulness curves for ViT-B/16 and DeiT-III-B/16; DAVE yields the most stable target-class prediction under progressive pixel removal.




Figure 5: Pixel Deletion results across ViT and B-cos-ViT models; DAVE demonstrates robust attribution faithfulness.
Quantitative and Qualitative Evaluation: Localization and Class Consistency
DAVE achieves substantial improvements (+1.79 to +11.92 percentage points Grid Pointing Game; up to +9.14 EnergyPG for B-cos models) across multiple ViT and B-cos architectures on ImageNet-1k.
Qualitative comparisons show that DAVE attributions are sharper, object-aligned, and with minimal background or grid noise compared to baseline methods, regardless of patch resolution or transformer-specific artifacts.

Figure 6: DAVE vs. prior attribution methods (SmoothGrad, IntGrad, AttnLRP, C-LRP, LeGrad) on ImageNet-1k; DAVE yields sharper, object-centric, and spatially coherent maps.
For inherently interpretable B-cos ViTs, DAVE improves upon the inherent explanations by further reducing background responses and localizing attribution sharply to object features.

Figure 7: Attribution maps for B-cos ViTs; DAVE delivers sharper, more object-aligned explanations than inherent B-cos explanations.

Figure 8: DAVE explanations highlighting object-centric features on B-cos ViTs, maintaining robustness to architectural variations such as the convolutional stem.

Figure 9: DAVE compared to post-hoc methods on DINO ViT-B/16; DAVE produces spatially coherent, object-aligned maps with suppressed patch artifacts.
DAVE also demonstrates class consistency: across images from the same class but with varying pose, background, or scale, DAVE highlights recurring, semantically meaningful cues (characteristic textures, contours, parts).

Figure 10: Examples of DAVE detecting class-consistent features across multiple validation images per ImageNet-1k class.
Robustness, Scalability, and Analysis
DAVE displays high robustness under input augmentations, input noise, and rotations. Rotational and noise sensitivity analyses reveal comparatively stable changes in model output and attribution maps. Convergence analysis verifies that DAVE attributions stabilize quickly (within 50 samples), ensuring computational efficiency.

Figure 11: Convergence analysis of DAVE attribution for multiple ViT variants, showing rapid stabilization as sample count increases.

Figure 12: Model output stability under input rotations, relevant to defining the spatial transformation neighborhood for DAVE.

Figure 13: Noise sensitivity analysis demonstrates stable DAVE attribution under additive Gaussian input perturbations.
Practical Implications and Theoretical Significance
DAVE provides a reliable post-hoc attribution mechanism for ViTs, suitable for high-stakes domains such as medical vision, autonomous vehicles, or auditing safety-critical systems. By removing architecture-induced artifacts and isolating semantically meaningful features, DAVE enables practitioners to audit model decisions, identify biases, and debug misclassifications in a manner closely aligned with visual evidence.
Theoretically, DAVE advances explainability by grounding attribution in operator decomposition and equivariant averaging, bridging mathematical rigor with practical visualization. DAVE’s structured approach is extensible to other architectures and pretraining modalities (e.g., self-supervised DINO models), and can improve attribution quality even for inherently interpretable networks.
Future Directions in Explainable AI
Opportunities for further research include optimizing DAVE's aggregation operator via adaptive, data-driven selection of transformation groups; efficient sampling schemes to reduce computational overhead; and deeper integration with multi-modal transformer architectures. Extending DAVE to spatiotemporal models or generative vision transformers could further enhance interpretability in complex visual domains.
Conclusion
DAVE represents a mathematically principled, architecture-aware attribution framework for ViTs, combining gradient decomposition, equivariant averaging, and low-pass filtering to deliver robust, fine-grained, and class-consistent attribution maps. Empirical evaluation demonstrates substantial gains in localization, faithfulness, and qualitative interpretability versus state-of-the-art baselines, across both conventional ViTs and inherently interpretable architectures. DAVE’s methodology offers a path toward reliable explainability for high-capacity vision models, facilitating model auditing, debugging, and safe deployment in sensitive applications.