Determine AURA’s robustness against adaptive adversarial attacks

Determine whether AURA remains reliable against adversarially crafted URLs designed to trigger early ham exits and semantically obfuscated phishing messages designed to evade the DistilBERT-LoRA encoder.

Background

The adaptive routing policy creates a potential attack surface: an attacker who understands the uncertainty thresholds could construct URLs that appear structurally benign and induce a low-uncertainty early ham decision, thereby bypassing semantic analysis.

A separate unresolved threat concerns phishing messages whose semantics are deliberately obfuscated to evade the Layer 2 DistilBERT-LoRA classifier. The paper identifies evaluation against both attack classes, together with adversarial training and certified robustness methods, as necessary directions for resolving this uncertainty.

References

AURA's routing component introduces a specific attack surface: an adversary aware of the routing policy could craft structurally benign-looking URLs to produce low uncertainty scores and trigger an early ham exit, bypassing DCAM entirely. Semantically obfuscated phishing messages engineered to evade the DistilBERT encoder represent a complementary threat to the Layer~2 component.

AURA: Adaptive Uncertainty-Routed Analysis for Email Threat Detection  (2609.19873 - Berjawi et al., 17 Sep 2026) in Section 7.1, “Adversarial Robustness”