---
title: 'DLFD: Distribution-Level Feature Distancing'
url: https://www.emergentmind.com/topics/distribution-level-feature-distancing-dlfd
type: topic
---

# DLFD: Distribution-Level Feature Distancing

Distribution-Level Feature Distancing (DLFD) is a machine unlearning methodology designed to address the fundamental trade-off between effective instance forgetting and the preservation of model utility. DLFD operates at the level of feature distributions, employing optimal transport metrics with entropic regularization in deep neural networks to synthesize data whose representations are explicitly distanced from the distribution of forget samples. This approach responds to the challenge of correlation collapse, whereby standard instance-level error maximization undermines class-separating features and degrades test performance during unlearning. DLFD achieves state-of-the-art results in utility-forgetting trade-offs on facial recognition tasks, validating its efficacy in scenarios demanding rigorous privacy guarantees and post hoc data removal [2409.14747].

## 1. Motivation and Background

The imperative for machine unlearning arises from the right to be forgotten in privacy-critical AI systems, such as facial recognition platforms subject to user-initiated data erasure. Traditional unlearning techniques either manipulate internal network parameters (e.g., via Fisher information regularization, teacher–student schemes) or synthesize adversarial perturbations at the instance level to induce forgetting. However, instance-level maximization of error on retained samples leads to correlation collapse—the phenomenon whereby task-discriminative feature correlations are weakened together with identity-specific features, resulting in degraded class separation in the neural feature space and substantial accuracy loss on retained tasks [2409.14747]. This scenario necessitates a training framework that enforces separation of forget and retain distributions without sacrificing task-relevant information.

## 2. Mathematical Framework and Core Mechanism

DLFD lifts the unlearning objective to the level of distributions in the model’s feature space. Consider feature vectors $F(x) \in \mathbb{R}^d$ from the penultimate layer of a deep network $\Theta$. For a batch of $n$ retained samples $(x_i, y_i)$ and $n$ forget samples $(x'_j, y'_j)$, DLFD defines the empirical distributions:
\[
\mu = \frac{1}{n} \sum_{i=1}^n \delta_{F(x_i)}, \qquad \nu = \frac{1}{n} \sum_{j=1}^n \delta_{F(x'_j)}.
\]
The optimal transport (OT) distance between $\mu$ and $\nu$ is
\[
\mathcal{D}(\mu, \nu) = \inf_{\gamma \in \Pi(\mu, \nu)} \mathbb{E}_{(w, w') \sim \gamma}[c(w, w')],
\]
where the cost $c(w, w') = 1 - \langle w, w'\rangle / \|w\| \|w'\|$ is the cosine distance, and $\Pi(\mu, \nu)$ denotes couplings with marginals $\mu$ and $\nu$. In practice, computation employs entropically regularized OT:
\[
T^\lambda = \arg\min_{T \in \Pi(\mu, \nu)} \langle T, C \rangle - \frac{1}{\lambda} \sum_{i,j} T_{ij} \log T_{ij},
\]
where $C_{ij} = c(F(x_i), F(x'_j))$ and $\lambda > 0$ controls regularization. The value $\ell_{OT} = \langle T^\lambda, C \rangle$ serves as a smooth, differentiable proxy for the separation between batch distributions.

Data synthesis perturbs each retained input $x_i$ via a single adversarial step:
\[
x_i^* \leftarrow x_i + \alpha \, \mathrm{sign} \big( \nabla_{x_i} [\ell_{OT} - \lambda_{CE} \ell_{CE}(y_i, \Theta(x_i))] \big),
\]
where $\ell_{CE}$ is the cross-entropy loss, $\alpha$ is a step size, and $\lambda_{CE} = \lambda (k/K)$ increases from $0$ to $1$ linearly over $K$ batch iterations, mediating between feature distancing and task fidelity. This maximizes the OT metric between retained and forget distributions while preserving class-specific feature alignment.

## 3. Algorithmic Procedure and Practical Usage

The full DLFD algorithm iterates over $K$ mini-batches in a single epoch:

1. Sample equal-sized batches of retained and forget data.
2. Compute the forgetting score via a membership-inference attack classifier. If above a predefined threshold, initialize $x_i^* \leftarrow x_i$ and apply $M$ inner gradient updates as above.
3. After $M$ updates, assemble the perturbed retained batch for unlearning and compute the batch cross-entropy training loss.
4. Update network parameters $\theta \leftarrow \theta - \gamma \nabla_\theta \ell_{train}$.

Because batch updates generate synthetic data with explicitly distanced feature distributions, one epoch typically suffices to reduce membership-inference attack accuracy on forget samples to near-random, while the gradual increase in $\lambda_{CE}$ preserves model utility by mitigating correlation collapse [2409.14747].

## 4. Empirical Results and Quantitative Evaluation

DLFD has been validated on facial recognition benchmarks under the instance-unlearning protocol:

- MUFAC age estimation: 8 classes, 10,025 training images, 1,500 forget
- RAF-DB emotion recognition: 7 classes, 11,044 training, 3,314 forget
- MUCAC multi-attribute classification: 3 binary labels, 25,933 training, 10,548 forget

Tested architectures include ResNet-18, DenseNet-121, and EfficientNet-B0. Compared methods encompass Retraining, Fine-tuning, NegGrad, CF-k, EU-k, UNSIR, BadTeaching, and SCRUB. Forgetting performance is measured via a membership-inference attack’s accuracy:
\[
\mathrm{ForgettingScore} = |{\mathrm{MIA}_{Acc}} - 0.5| \times 2\quad(\text{ideal} = 0),
\]
with utility given by test set accuracy, and overall efficacy summarized by the NoMUS score:
\[
\mathrm{NoMUS} = \frac{1}{2} \big(P(\hat{y} = y) + 1 - \mathrm{ForgettingScore}\big).
\]

DLFD attains the most favorable trade-off. For instance, in ResNet-18 age classification, DLFD yields test accuracy $0.6166$ (vs. $0.6329$ pre-unlearning), a forgetting score $0.0385$ (vs. $0.1923$ baseline), and NoMUS $0.7698$ (highest among compared methods). Similar superlative results are recorded for emotion (NoMUS $0.7934$), multi-attribute ($0.9283$), and gender ($0.9192$) tasks. t-SNE visualizations confirm that DLFD preserves class-separating structure otherwise destroyed by error-based perturbation. Loss histograms across unseen and forget samples indicate that DLFD aligns with retraining, exhibiting randomized MIA predictions on forget data [2409.14747].

## 5. Theoretical Motivation and Design Intuitions

DLFD's efficacy is grounded in two core rationales. First, the optimal transport (OT) metric reflects the global structure of high-dimensional feature manifolds, guaranteeing that the separation between retained and forgotten data arises at the distributional rather than instance level. Second, the entropic regularization of the transport plan ensures smooth updates, thereby avoiding the creation of boundary outliers that would otherwise impair classification performance. The one-step adversarial update with dynamically calibrated $\lambda_{CE}$ maintains a balance between pushing retained features away from forget regions and retaining essential feature-label correlations, directly addressing the correlation collapse observed in prior works [2409.14747].

## 6. Limitations and Future Directions

DLFD is subject to several practical and theoretical limitations. Its effectiveness may diminish when feature distributions of retained and forgotten samples are highly overlapping. The method employs membership-inference as the primary forgetting metric, potentially restricting its guarantees under adversarial attack models not based on MIA. Scalability is feasible for moderate-scale convolutional architectures, but extension to large-scale models may necessitate approximate OT solutions or alternative feature projections. Areas for future work include incorporating additional distributional metrics (e.g., sliced-Wasserstein), adaptive or schedule-based tuning of $\lambda$, multi-epoch variants for finer-grained unlearning, and exploration of alternative unlearning performance metrics [2409.14747].

## 7. Relationship to Other Distributional Metrics

DLFD’s distribution-level approach is conceptually linked to recent advances in statistical distances such as energy distance for measuring feature heterogeneity in distributed learning systems. Notably, alternatives like the normalized energy coefficient $H(X,Y)$ provide metric-based distributional comparisons sensitive to location, scale, and higher moments, while supporting low-complexity Taylor-approximate computation for large datasets [2501.16174]. While DLFD focuses on OT-based cosine cost to dissociate forget and retain distributions within a single model, energy distance has been adopted for weighting penalties in federated optimization to counteract distributional drift. This suggests potential opportunities for cross-pollination between the DLFD paradigm and federated settings facing non-IID data distributions.

Source: https://www.emergentmind.com/topics/distribution-level-feature-distancing-dlfd