Nyquist-limited overcomplete FouT basis representation

Determine whether the number of basis vectors in a FouT state-space model can exceed the number of sample points in its memory window without encoding frequencies above the Nyquist frequency, thereby increasing frequency resolution below the PCM Nyquist frequency and potentially improving speech-enhancement performance.

Background

The paper uses FouT and LegT state-space model initializations to encode speech into coefficient sequences. The speech-enhancement experiments indicate that increasing the state dimension relative to the number of samples in the memory window can cause basis functions to represent frequencies above the PCM Nyquist frequency, which is problematic when processing PDM signals because high-frequency quantization noise may be aliased into the speech band during reconstruction.

The authors explicitly leave unresolved whether FouT basis vectors can be focused on frequencies below the PCM Nyquist frequency while allowing more basis vectors than memory-window sample points. They compare the desired construction to zero-padding in short-time Fourier transform processing, where padding increases frequency resolution without extending the represented frequency band.

References

Nevertheless, our speech enhancement experiments raised a critical question: is it possible to increase the number of basis vectors in the SSM such that there are more basis vectors than sample points in the memory window without encoding frequencies higher than the Nyquist frequency? This would be analoguous to standard STFT processing where the STFT can be padded to increase frequency resolution. Therefore, future work will investigate how the FouT basis vectors could be focused to represent the frequencies under the PCM Nyquist frequency, analogous to the STFT, thereby increasing the frequency resolution of the FouT SSM under the PCM Nyquist frequency and possibly leading to better speech enhancement performance.

— Multirate State Space Models for End-to-End Processing of Pulse Density Modulated Speech Signals  (2608.28472 - Boulanger et al., 28 Aug 2026) in Section 6, Conclusion