General nonlinear concept erasure

Establish effective general nonlinear concept-erasure methods that remove concept-specific information while avoiding the utility and computational limitations of existing approaches.

Background

The paper distinguishes linear guardedness and linear erasure methods, such as INLP, RLACE, and LEACE, from the more difficult objective of general nonlinear concept erasure. Existing nonlinear approaches—including adversarial training, information-theoretic optimization, Bayesian optimization, filtering, and density matching—are described as often sacrificing downstream utility and incurring substantial computational costs. The authors therefore identify general nonlinear concept erasure as an unresolved challenge, while the proposed MUtE framework addresses a particular continuous, discrete-concept setting through conditional density matching and a translational geometric bias.

References

However, general (non-linear) concept erasure remains an open challenge.

MUtE: A Dual Framework for Concept Erasure and Counterfactual Interventions  (2609.11253 - Saillenfest, 10 Sep 2026) in Section 1, Introduction