Mechanism underlying distillation’s fairness improvement
Determine whether teacher pseudo-label averaging over the full input distribution or smoothing of the student’s decision boundary causes knowledge distillation to narrow demographic fairness gaps in Whisper automatic speech recognition models.
References
Why distillation narrows where pruning compounds is open. Training on teacher pseudo-labels averages the student's loss over the teacher's full input distribution, which may regularize against the worst-group behavior the teacher already exhibits, and may also push the student toward a smoother decision boundary that is less peaked on majority-group features. Adjudicating between these would require training-data attribution or controlled pseudo-label experiments, which we leave to future work.