Make acoustics necessary in latent acoustic-to-articulatory prediction

Develop an acoustic-to-articulatory latent-rollout model in which future vocal-tract representations are causally driven by acoustic conditioning rather than being predictable from the initial visual frames alone.

Background

The paper extends the frozen Arti-JEPA encoder with an acoustic-to-articulatory inversion task based on latent rollout: a predictor receives several seed vocal-tract frames and a frozen audio embedding, then predicts future vocal-tract latents. Across several conditioning mechanisms, shuffling the audio changes rollout error by only a few percent of the error that acoustic conditioning could theoretically remove.

The authors attribute this failure to the predictability of vocal-tract motion over the 0.64-second prediction window: the seed frames already determine most of the future dynamics, allowing the predictor to ignore the audio. They identify longer sequence prediction, residual prediction against an audio-free predictor, and discriminative articulatory objectives as possible ways to make acoustic information necessary for the task.

References

Making the acoustics necessary rather than merely available is the open problem here (Appendix ~\ref{sec:aai}).

Arti-JEPA: Adapting Video World Model to Real-Time MRI of the Vocal Tract for Speech-Production Analysis  (2609.09757 - Nguyen et al., 9 Sep 2026) in Section 5, “Limitations and discussion,” paragraph “Conditioned prediction: a negative result”; Appendix, Section “Acoustically-conditioned latent rollout: an initial attempt”