Cross-family generalization of pruning and INT4 fairness effects

Determine whether the demographic fairness effects of Whisper pruning and sub-8-bit quantization, particularly Wanda pruning and INT4 quantization, generalize to speech-recognition architectures outside the Whisper family.

Background

The principal experiments are restricted to the Whisper family to isolate compression effects from differences in architecture, tokenizer, training data, and audio encoder. An appendix probe on IBM Granite-4.0-1b-speech evaluates only the FP16-to-INT8 comparison.

Because pruning and INT4 quantization were not tested on Granite, the probe does not resolve whether the major fairness findings extend beyond Whisper. The paper explicitly identifies cross-family experiments on these compression axes as future work.

References

The probe does not adjudicate whether the pruning or INT4 findings generalize; cross-family experiments on those axes are left to future work.

— Temporal Taxation Compounds Under Post-Training Compression of Whisper Models  (2609.28739 - Ginjala et al., 23 Sep 2026) in Section 6, “Concluding Remarks and Discussion”; Section “Granite-4.0 generalization probe”