Posterior-head adaptation under domain shift

Develop a method for adapting the SPSI posterior head under domain shift without collapsing the estimated soft speaker posterior \(\hat{\mathbf{P}}\), thereby improving transfer performance beyond the current joint-update procedure.

Background

Soft Posterior Speaker Injection (SPSI) estimates frame-level soft speaker shares with a posterior head and injects them into Whisper through encoder FiLM layers and decoder speaker-memory prompts. For LibriCSS transfer, freezing the posterior head while adapting the remaining model improves performance, whereas jointly updating the posterior head currently harms transfer. The paper therefore identifies calibration-preserving adaptation of the posterior head under domain shift as an unresolved methodological direction.

References

A remaining issue is posterior calibration under domain shift: jointly updating $\phi$ currently hurts transfer, so adapting the head without collapsing $\hat{\mathbf{P}}$ is an open direction.

Soft Posterior Speaker Injection for Multi-Talker Speech Recognition  (2609.01287 - Zhu et al., 1 Sep 2026) in Section Conclusion and Future Work