Close the generalization gap against encoding-level adversarial attacks

Close the remaining robustness gap of minimally adversarially trained BERT against mechanistically distinct, encoding-level adversarial attacks, particularly homoglyph substitution, by augmenting training with at least one encoding-level attack and evaluating whether this resolves the observed failure to generalize uniformly across perturbation classes.

Background

The paper evaluates a BERT spam detector trained on a limited set of adversarial examples and tests it against held-out attacks. Robustness transfers nearly completely to held-out attacks within the same transformation class, but transfer to structurally different Unicode encoding and rendering attacks is weaker.

Homoglyph substitution is identified as the clearest unresolved case: minimally adversarially trained BERT improves substantially over the non-adversarially trained model but remains 19.8 percentage points below the fully adversarially trained model. The authors attribute this gap to the cross-script Unicode mechanism and state that closing it requires adding an encoding-level attack to training; the problem is explicitly left for future work.

References

Closing this specific gap would require augmenting training with at least one encoding-level attack, which we leave to future work alongside the broader ensemble defenses evaluated in the next section.

Johnny Still Receives Spam SMS: Assessing the Robustness of SMS Spam Detection  (2609.01171 - Salman et al., 1 Sep 2026) in Section 8, subsection “Can BERT Withstand the Unseen?: Exploring General Robustness”