---
title: 'FERD: Fairness-Enhanced Robustness Distillation'
url: https://www.emergentmind.com/topics/fairness-enhanced-data-free-robustness-distillation-ferd
type: topic
---

# FERD: Fairness-Enhanced Robustness Distillation

Searching arXiv for the specified paper to ground the article in the current record.
Fairness-Enhanced Data-Free Robustness Distillation (FERD) is a framework for Data-Free Robustness Distillation (DFRD) that aims to transfer adversarial robustness from a robust teacher model to a lightweight student without access to the teacher’s original training data, while explicitly addressing robust fairness across categories [2509.20793]. In DFRD, a generator synthesizes samples for distillation; FERD modifies this paradigm by adjusting both the proportion and the distribution of adversarial examples. It is presented as the first fairness-enhanced data-free robustness distillation framework, motivated by two observed problems in existing methods: severe disparity of robustness across classes and instability of the student’s robustness across attack targets [2509.20793].

## 1. Problem formulation and motivating failure modes

In the FERD setting, the goal is to transfer adversarial robustness from a teacher \(f^T\) to a student \(f^S\) without any access to the training data. The basic DFRD approach is to train a generator \(G\) to synthesize “fake” samples \(\{(x_i,y_i)\}\), and then use both clean and adversarially perturbed versions of these samples for distillation [2509.20793].

FERD is motivated by two fairness-related deficiencies in existing DFRD. First, even when synthetic samples are drawn with equal class proportions, the student’s per-class robustness under attack often shows large gaps. Some classes rapidly inherit robustness, whereas others remain comparatively vulnerable. Second, standard adversarial-example generation tends to push samples toward the most confusable target class, so the student can become specialized in defending against only a small subset of target directions. The reported consequence is poor defense when the attacker chooses other target classes [2509.20793].

The framework therefore addresses fairness in two distinct senses. One concerns disparity across source categories, where some classes receive weaker robustness transfer than others. The other concerns imbalance across attack targets, where the induced adversarial distribution is concentrated on a narrow set of vulnerable directions. This suggests that FERD treats robust fairness not merely as class-balanced sampling, but as a joint problem of sample allocation and attack-direction diversity.

## 2. Robustness-guided class reweighting

FERD’s first component is a robustness-guided class reweighting strategy designed to allocate more synthetic samples to weakly robust categories. The procedure begins by generating adversarial examples \(x_i^{adv}\) from each synthetic sample \(x_i\), for example with PGD-20. For each such example, the teacher’s adversarial margin is computed as

\[
m_i \;=\;\bigl(f^T(x_i^{adv})\bigr)_{\,y_i}\;-\;\max_{j\neq y_i}\bigl(f^T(x_i^{adv})\bigr)_j.
\]

A smaller, or negative, margin indicates greater vulnerability. Class-level vulnerability is then defined by averaging the negative margins within each class:

\[
D_c \;=\;\frac{1}{N_c}\sum_{i:\,y_i=c}(-m_i),
\]

where \(N_c\) is the number of synthesized samples with label \(c\). These vulnerability scores are transformed into a sampling distribution through a softmax,

\[
w_c \;=\;\frac{\exp(D_c)}{\sum_{c'}\exp(D_{c'})},
\]

and subsequent generator updates sample labels from \(y\sim \mathrm{Categorical}(\{w_c\})\) rather than from a uniform distribution [2509.20793].

The operational effect is that \(G\) synthesizes more samples for classes with high vulnerability \(D_c\). In the formulation given for FERD, this component directly targets the observed disparity in per-class robustness. The paper further states that reweighting compensates weak classes by supplying more training samples and thereby closes the inter-class robustness gap, empirically reducing the normalized standard deviation of per-class robustness [2509.20793].

## 3. Fairness-Aware Example generation

FERD’s second component is Fairness-Aware Example (FAE) generation. Its stated goal is to produce benign samples whose non-robust features do not strongly bias toward any single class. The construction begins by identifying non-robust feature channels using an information-bottleneck style scheme at an intermediate layer \(\ell\) of the teacher network [2509.20793].

An intermediate representation is extracted as \(Z=f^T_\ell(x)\). Gaussian noise scaled by a learnable vector \(\lambda_r\) is then injected:

\[
Z_I \;=\;f^T_\ell(x)\;+\;\mathrm{softplus}(\lambda_r)\,\epsilon,\quad
\epsilon\sim\mathcal{N}(0,I).
\]

The vector \(\lambda_r\) is optimized to preserve predictiveness of \(Z_I\) for \(y\) under the teacher while limiting information flow:

\[
\min_{\lambda_r}\; \underbrace{CE\bigl(f^T_{\ell+}(Z_I),y\bigr)}_{\text{prediction loss}}
\;+\;\beta\sum_{c}\Bigl(\tfrac{v_c}{\lambda_c^2}+\log\tfrac{\lambda_c^2}{v_c}-1\Bigr),
\]

where \(v_c=\mathrm{Var}(Z_I^{(c)})\) is the variance of channel \(c\). A channel \(k\) is marked as non-robust if

\[
\lambda_k^2 < \max_{c'}\,\mathrm{Var}\bigl(Z_I^{(c')}\bigr),
\]

and the non-robust feature \(Z_{nr}\) is then formed by channel-wise masking [2509.20793].

FERD next imposes a uniformity constraint on the teacher’s output when only the non-robust features are provided:

\[
\mathcal{L}_{uni}
\;=\;KL\Bigl(\,\mathcal{U}\;\Big\|\;f^T_{\ell+}(Z_{nr})\Bigr),
\]

where \(\mathcal{U}\) denotes the uniform distribution over the \(C\) classes. The full generator loss for synthesizing FAEs \(x_F\) is

\[
\mathcal{L}_{gen}
=\lambda_{adv}\,KL\bigl(f^T(x_F),f^S(x_F)\bigr)
+\lambda_{bn}\,\mathcal{L}_{bn}
+\lambda_{oh}\,CE\bigl(f^T(x_F),y\bigr)
+\lambda_{uni}\,\mathcal{L}_{uni}.
\]

Here, \(\mathcal{L}_{bn}\) matches BatchNorm statistics of teacher and student, while the \(CE\) and KL terms enforce diversity and correctness [2509.20793].

The stated interpretation is that FAEs suppress the dominance of class-specific non-robust features and provide a more balanced representation across all categories. A plausible implication is that FERD reorients the generator away from artifacts that are predictive yet unevenly distributed across classes, which in turn is intended to improve worst-class robustness rather than only average robustness.

## 4. Uniform-Target Adversarial Example construction

FERD’s third component is the construction of Uniform-Target Adversarial Examples (UTAEs). The purpose is to diversify attack targets so that the student learns to defend against perturbations pushing toward every class equally, rather than predominantly toward the most vulnerable or most confusable classes [2509.20793].

Starting from a generated FAE \(x_F\), FERD applies a modified PGD procedure:

\[
x^{t+1}_U
\;=\;
\Pi_{x_F+\mathcal{S}\Bigl(
x^t_U
+\alpha\;\mathrm{sign}\bigl[\nabla_{x^t_U}\bigl(
KL(f^T(x_F),f^T(x^t_U))
-\gamma\,KL(\mathcal{U},f^T(x^t_U))
\bigr)\bigr]
\Bigr),
\]

where \(\Pi\) projects back into the \(\ell_\infty\) ball \(\mathcal{S}\). In the description of the method, the first KL term pulls the adversarial example away from the original FAE label, while the second term, weighted by \(\gamma\), pushes the attack toward a uniform target distribution. Repeating the update for \(T\) steps yields the UTAE \(x_U\) [2509.20793].

The student is then distilled with both FAEs and UTAEs according to

\[
\mathcal{L}_{stu}
=\lambda_1\,KL\bigl(f^T(x_F),f^S(x_F)\bigr)
+\lambda_2\,KL\bigl(f^T(x_F),f^S(x_U)\bigr).
\]

In the paper’s own justification, UTAE diversification ensures that the student is exposed to adversarial shifts toward all possible target classes, thereby reducing vulnerability to any single target. This indicates that FERD treats robustness transfer as sensitive not only to perturbation magnitude and source label, but also to the geometry of target-class allocation in adversarial space [2509.20793].

## 5. End-to-end algorithmic structure and rationale

The overall FERD algorithm takes as input a teacher \(f^T\), a student \(f^S\), a generator \(G\), and a number of epochs \(E\), together with the hyperparameters \(\lambda_{adv},\lambda_{bn},\lambda_{oh},\lambda_{uni},\gamma,\lambda_1,\lambda_2\). For each epoch, the procedure is described as follows: sample a batch of class labels according to the reweighting distribution \(\{w_c\}\); generate benign FAEs \(x_F^i=G(z_i,y_i)\); update \(\lambda_r\) and compute \(\mathcal{L}_{uni}\); update the generator by minimizing \(\mathcal{L}_{gen}\); generate UTAEs from each FAE via the uniform-target PGD; compute \(\mathcal{L}_{stu}\) and update the student; and periodically re-estimate margins \(\{m_i\}\) and recompute \(w_c\) [2509.20793].

The rationale offered for this design is explicit. Reweighting addresses weak classes by increasing their synthetic-sample allocation. FAEs suppress class-specific non-robust features, preventing the generator from over-emphasizing a few easy classes. UTAE diversification broadens the set of target directions encountered during robustness distillation. The paper additionally reports empirical bounds in Appendix A and B indicating that both uniform class sampling and uniform target distribution tighten the worst-case generalization bound under adversarial shifts [2509.20793].

Taken together, these components define a coupled mechanism in which class-level vulnerability estimates influence sample generation, feature-level uniformity reshapes benign synthetic data, and target-level uniformity reshapes adversarial synthetic data. This suggests that FERD frames fairness-enhanced robustness distillation as a multi-level balancing problem spanning class frequency, feature salience, and attack-direction coverage.

## 6. Experimental setting, reported results, and implementation details

FERD is evaluated on three public datasets: CIFAR-10, CIFAR-100, and Tiny-ImageNet. The teacher architectures are WideResNet-34-10 for CIFAR-10 and CIFAR-100, and PreActResNet-34 for Tiny-ImageNet. The student architectures are ResNet-18 and MobileNet-V2 for CIFAR, and PreActResNet-18 for Tiny-ImageNet. The attacks used for evaluation are FGSM, PGD-20, \(CW_\infty\), and AutoAttack (AA) [2509.20793].

The paper reports that FERD achieves state-of-the-art worst-class robustness under all adversarial attack settings. For MobileNet-V2 on CIFAR-10, the key quantitative gains in worst-class robustness are reported as \(15.1\%\) under FGSM, \(2.7\%\) under PGD-20, \(3.4\%\) under \(CW_\infty\), and \(6.4\%\) under AutoAttack. It is also stated that FERD yields the lowest Normalized Standard Deviation (NSD) in most cases, which is presented as signifying higher robust fairness [2509.20793].

The implementation details provided are specific. The generator is optimized with Adam using learning rate \(2\times 10^{-3}\), \(\beta_1=0.5\), and \(\beta_2=0.999\). The student is optimized with SGD using initial learning rate \(0.1\), momentum \(0.9\), and weight decay \(5\times 10^{-4}\). Distillation runs for 220 epochs, with 400 iterations per epoch for both \(G\) and \(f^S\). The batch size is 256 on CIFAR-10 and 512 on CIFAR-100 and Tiny-ImageNet. The feature extraction layer \(\ell\) is the last convolutional block of \(f^T\). The reported hyperparameters are \(\lambda_{adv}=1\), \(\lambda_{bn}=5\), \(\lambda_{oh}=1\), \(\lambda_{uni}=5\), \(\gamma=0.5\), \(\lambda_1=5/6\), and \(\lambda_2=1/6\) [2509.20793].

A common misconception in robustness distillation is that balanced class counts alone suffice to ensure balanced robustness outcomes. FERD is explicitly motivated by the observation that equal class proportion data can still produce substantial differences across categories, and that target imbalance during attack generation constitutes a separate failure mode. Within the scope of the reported experiments, the framework is therefore positioned not merely as a stronger DFRD variant, but as a method that reformulates robustness transfer in terms of worst-class robustness and robust fairness rather than overall robustness alone [2509.20793].

Source: https://www.emergentmind.com/topics/fairness-enhanced-data-free-robustness-distillation-ferd