Close the generalization gap against encoding-level adversarial attacks
Close the remaining robustness gap of minimally adversarially trained BERT against mechanistically distinct, encoding-level adversarial attacks, particularly homoglyph substitution, by augmenting training with at least one encoding-level attack and evaluating whether this resolves the observed failure to generalize uniformly across perturbation classes.
References
Closing this specific gap would require augmenting training with at least one encoding-level attack, which we leave to future work alongside the broader ensemble defenses evaluated in the next section.
— Johnny Still Receives Spam SMS: Assessing the Robustness of SMS Spam Detection
(2609.01171 - Salman et al., 1 Sep 2026) in Section 8, subsection “Can BERT Withstand the Unseen?: Exploring General Robustness”