Training target models for intent-consistent safety under surface variation
Determine how to train the target language model itself so that surface variants expressing the same underlying intent receive consistent, directionally separated refusal or compliance decisions.
References
These approaches motivate intent-aware and robustness-oriented safety training, but leave open how to train the target model itself so that same-intent surface variants receive consistent, directionally separated refusal/compliance decisions.
Finally, developing unlearning methods that explicitly optimize for contextual separation, rather than treating forgetting and retention independently, remains a key open challenge.
Regardless of which hypothesis holds, it remains an open question whether coherent alignment of LLM-based systems is realizable in practice.