Evaluate the effectiveness of machine unlearning in large language models
Determine a rigorous, standardized methodology to evaluate the effectiveness of machine unlearning in large language models, including clear criteria and metrics for assessing whether targeted knowledge has been removed rather than merely suppressed.
References
Unlearning in LLMs is crucial for managing sensitive data and correcting misinformation, yet evaluating its effectiveness remains an open problem.
Most practical unlearning setups, including full fine-tuning of a LLM, do not hold earlier layers fixed the way we do here, raising a question our results do not answer: why would suppression neurons appear in practice \citep{yang2026erase} if genuine deletion is available whenever upstream parameters are free to move? We do not know, but Appendix~\ref{app:encoder_freeze} discusses two hypotheses for why deletion may be less available in practice than this comparison implies.
We do not know whether the mirror/shadow distinction, or anything resembling it, is visible when the same excision procedure is applied to a standard unlearning benchmark such as TOFU \citep{maini2024tofu} or Bias in Bios \citep{dearteaga2019bias}, where the relevant features are not hand-constructed to be antipodal and may not be separable from the rest of the representation in as clean a way.
Moreover, the training data cutoffs of different models are not identical. The relative timing between API deprecation and a model's training data cutoff may therefore affect whether the deprecated API was actually exposed to the model during pre-training. Although we further analyze the influence of this factor on unlearning performance, it remains difficult to determine the extent to which specific API knowledge was present in the actual training corpora.