Mechanism underlying distillation’s fairness improvement

Determine whether teacher pseudo-label averaging over the full input distribution or smoothing of the student’s decision boundary causes knowledge distillation to narrow demographic fairness gaps in Whisper automatic speech recognition models.

Background

The experiments show that distillation narrows demographic fairness gaps in most evaluated teacher–student, precision, and dataset combinations, whereas pruning often compounds those gaps. The paper proposes two possible explanations: pseudo-label training may regularize worst-group behavior, or it may produce a smoother decision boundary that is less dependent on majority-group features.

The authors state that distinguishing these explanations would require training-data attribution or controlled pseudo-label experiments, neither of which is performed in the study.

References

Why distillation narrows where pruning compounds is open. Training on teacher pseudo-labels averages the student's loss over the teacher's full input distribution, which may regularize against the worst-group behavior the teacher already exhibits, and may also push the student toward a smoother decision boundary that is less peaked on majority-group features. Adjudicating between these would require training-data attribution or controlled pseudo-label experiments, which we leave to future work.

— Temporal Taxation Compounds Under Post-Training Compression of Whisper Models  (2609.28739 - Ginjala et al., 23 Sep 2026) in Section 5, paragraph “Candidate mechanisms”