DLFD: Distribution-Level Feature Distancing
- DLFD is a machine unlearning methodology that leverages optimal transport metrics and entropic regularization to separate retained and forgotten feature distributions.
- It addresses correlation collapse by synthesizing perturbed data that preserves class-specific features while enforcing privacy through instance forgetting.
- Empirical results on facial recognition benchmarks confirm that DLFD maintains high model utility and significantly reduces membership-inference attack accuracy.
Distribution-Level Feature Distancing (DLFD) is a machine unlearning methodology designed to address the fundamental trade-off between effective instance forgetting and the preservation of model utility. DLFD operates at the level of feature distributions, employing optimal transport metrics with entropic regularization in deep neural networks to synthesize data whose representations are explicitly distanced from the distribution of forget samples. This approach responds to the challenge of correlation collapse, whereby standard instance-level error maximization undermines class-separating features and degrades test performance during unlearning. DLFD achieves state-of-the-art results in utility-forgetting trade-offs on facial recognition tasks, validating its efficacy in scenarios demanding rigorous privacy guarantees and post hoc data removal (Choi et al., 2024).
1. Motivation and Background
The imperative for machine unlearning arises from the right to be forgotten in privacy-critical AI systems, such as facial recognition platforms subject to user-initiated data erasure. Traditional unlearning techniques either manipulate internal network parameters (e.g., via Fisher information regularization, teacher–student schemes) or synthesize adversarial perturbations at the instance level to induce forgetting. However, instance-level maximization of error on retained samples leads to correlation collapse—the phenomenon whereby task-discriminative feature correlations are weakened together with identity-specific features, resulting in degraded class separation in the neural feature space and substantial accuracy loss on retained tasks (Choi et al., 2024). This scenario necessitates a training framework that enforces separation of forget and retain distributions without sacrificing task-relevant information.
2. Mathematical Framework and Core Mechanism
DLFD lifts the unlearning objective to the level of distributions in the model’s feature space. Consider feature vectors from the penultimate layer of a deep network . For a batch of retained samples and forget samples , DLFD defines the empirical distributions: The optimal transport (OT) distance between and is
where the cost 0 is the cosine distance, and 1 denotes couplings with marginals 2 and 3. In practice, computation employs entropically regularized OT: 4 where 5 and 6 controls regularization. The value 7 serves as a smooth, differentiable proxy for the separation between batch distributions.
Data synthesis perturbs each retained input 8 via a single adversarial step: 9 where 0 is the cross-entropy loss, 1 is a step size, and 2 increases from 3 to 4 linearly over 5 batch iterations, mediating between feature distancing and task fidelity. This maximizes the OT metric between retained and forget distributions while preserving class-specific feature alignment.
3. Algorithmic Procedure and Practical Usage
The full DLFD algorithm iterates over 6 mini-batches in a single epoch:
- Sample equal-sized batches of retained and forget data.
- Compute the forgetting score via a membership-inference attack classifier. If above a predefined threshold, initialize 7 and apply 8 inner gradient updates as above.
- After 9 updates, assemble the perturbed retained batch for unlearning and compute the batch cross-entropy training loss.
- Update network parameters 0.
Because batch updates generate synthetic data with explicitly distanced feature distributions, one epoch typically suffices to reduce membership-inference attack accuracy on forget samples to near-random, while the gradual increase in 1 preserves model utility by mitigating correlation collapse (Choi et al., 2024).
4. Empirical Results and Quantitative Evaluation
DLFD has been validated on facial recognition benchmarks under the instance-unlearning protocol:
- MUFAC age estimation: 8 classes, 10,025 training images, 1,500 forget
- RAF-DB emotion recognition: 7 classes, 11,044 training, 3,314 forget
- MUCAC multi-attribute classification: 3 binary labels, 25,933 training, 10,548 forget
Tested architectures include ResNet-18, DenseNet-121, and EfficientNet-B0. Compared methods encompass Retraining, Fine-tuning, NegGrad, CF-k, EU-k, UNSIR, BadTeaching, and SCRUB. Forgetting performance is measured via a membership-inference attack’s accuracy: 2 with utility given by test set accuracy, and overall efficacy summarized by the NoMUS score: 3
DLFD attains the most favorable trade-off. For instance, in ResNet-18 age classification, DLFD yields test accuracy 4 (vs. 5 pre-unlearning), a forgetting score 6 (vs. 7 baseline), and NoMUS 8 (highest among compared methods). Similar superlative results are recorded for emotion (NoMUS 9), multi-attribute (0), and gender (1) tasks. t-SNE visualizations confirm that DLFD preserves class-separating structure otherwise destroyed by error-based perturbation. Loss histograms across unseen and forget samples indicate that DLFD aligns with retraining, exhibiting randomized MIA predictions on forget data (Choi et al., 2024).
5. Theoretical Motivation and Design Intuitions
DLFD's efficacy is grounded in two core rationales. First, the optimal transport (OT) metric reflects the global structure of high-dimensional feature manifolds, guaranteeing that the separation between retained and forgotten data arises at the distributional rather than instance level. Second, the entropic regularization of the transport plan ensures smooth updates, thereby avoiding the creation of boundary outliers that would otherwise impair classification performance. The one-step adversarial update with dynamically calibrated 2 maintains a balance between pushing retained features away from forget regions and retaining essential feature-label correlations, directly addressing the correlation collapse observed in prior works (Choi et al., 2024).
6. Limitations and Future Directions
DLFD is subject to several practical and theoretical limitations. Its effectiveness may diminish when feature distributions of retained and forgotten samples are highly overlapping. The method employs membership-inference as the primary forgetting metric, potentially restricting its guarantees under adversarial attack models not based on MIA. Scalability is feasible for moderate-scale convolutional architectures, but extension to large-scale models may necessitate approximate OT solutions or alternative feature projections. Areas for future work include incorporating additional distributional metrics (e.g., sliced-Wasserstein), adaptive or schedule-based tuning of 3, multi-epoch variants for finer-grained unlearning, and exploration of alternative unlearning performance metrics (Choi et al., 2024).
7. Relationship to Other Distributional Metrics
DLFD’s distribution-level approach is conceptually linked to recent advances in statistical distances such as energy distance for measuring feature heterogeneity in distributed learning systems. Notably, alternatives like the normalized energy coefficient 4 provide metric-based distributional comparisons sensitive to location, scale, and higher moments, while supporting low-complexity Taylor-approximate computation for large datasets (Fan et al., 27 Jan 2025). While DLFD focuses on OT-based cosine cost to dissociate forget and retain distributions within a single model, energy distance has been adopted for weighting penalties in federated optimization to counteract distributional drift. This suggests potential opportunities for cross-pollination between the DLFD paradigm and federated settings facing non-IID data distributions.