Sensitivity of adversarial effects to perturbation type

Investigate how the results of the paired HotFlip-versus-random adversarial framework vary across other forms of substitution perturbations.

Background

The paper compares gradient-guided HotFlip token substitutions with rate- and position-matched random token substitutions and finds stronger behavioral and representational disruption for HotFlip. However, the adversarial comparison is limited primarily to token substitutions. It remains unresolved whether the observed adversarial effect is specific to this perturbation type or generalizes to other substitution procedures.

References

Additionally, the paired adversarial framework in RQ4 can be extended to other forms of substitution perturbations, as it is unclear how the adversarial result may be sensitive to perturbation type.

How Perturbations Propagate: A Multi-Level Analysis of Robustness in Large Language Models  (2609.03322 - Chan et al., 3 Sep 2026) in Conclusion