Evaluate the effectiveness of machine unlearning in large language models

Determine a rigorous, standardized methodology to evaluate the effectiveness of machine unlearning in large language models, including clear criteria and metrics for assessing whether targeted knowledge has been removed rather than merely suppressed.

Background

The paper introduces the Stimulus-Knowledge Entanglement-Behavior (SKeB) framework to study how persuasive prompting and knowledge entanglement influence residual recall in unlearned LLMs. Despite proposing SKeB, the authors explicitly note that reliably evaluating whether unlearning has truly removed specific information remains unresolved.

This open problem is foundational to privacy, safety, and compliance claims around unlearning, since current approaches may suppress direct recall while leaving indirect retrieval pathways intact via framing or entanglement.

References

Unlearning in LLMs is crucial for managing sensitive data and correcting misinformation, yet evaluating its effectiveness remains an open problem.

Most practical unlearning setups, including full fine-tuning of a LLM, do not hold earlier layers fixed the way we do here, raising a question our results do not answer: why would suppression neurons appear in practice \citep{yang2026erase} if genuine deletion is available whenever upstream parameters are free to move? We do not know, but Appendix~\ref{app:encoder_freeze} discusses two hypotheses for why deletion may be less available in practice than this comparison implies.

— Hidden not Deleted: How Networks Suppress Entangled Features  (2609.27593 - Samanta et al., 23 Sep 2026) in Section 7, paragraph “Limitations”

We do not know whether the mirror/shadow distinction, or anything resembling it, is visible when the same excision procedure is applied to a standard unlearning benchmark such as TOFU \citep{maini2024tofu} or Bias in Bios \citep{dearteaga2019bias}, where the relevant features are not hand-constructed to be antipodal and may not be separable from the rest of the representation in as clean a way.

— Hidden not Deleted: How Networks Suppress Entangled Features  (2609.27593 - Samanta et al., 23 Sep 2026) in Appendix, Section “Directions for Empirical Validation,” paragraph “Validation on standard benchmarks”

Moreover, the training data cutoffs of different models are not identical. The relative timing between API deprecation and a model's training data cutoff may therefore affect whether the deprecated API was actually exposed to the model during pre-training. Although we further analyze the influence of this factor on unlearning performance, it remains difficult to determine the extent to which specific API knowledge was present in the actual training corpora.

— What Was Once Learned May Need to Be Unlearned: Machine Unlearning for Deprecated API Knowledge in Large Language Models  (2609.25786 - Liu et al., 22 Sep 2026) in Section 7, Discussion, subsection “Threats to Validity,” item (4)