Make acoustics necessary in latent acoustic-to-articulatory prediction
Develop an acoustic-to-articulatory latent-rollout model in which future vocal-tract representations are causally driven by acoustic conditioning rather than being predictable from the initial visual frames alone.
References
Making the acoustics necessary rather than merely available is the open problem here (Appendix ~\ref{sec:aai}).
— Arti-JEPA: Adapting Video World Model to Real-Time MRI of the Vocal Tract for Speech-Production Analysis
(2609.09757 - Nguyen et al., 9 Sep 2026) in Section 5, “Limitations and discussion,” paragraph “Conditioned prediction: a negative result”; Appendix, Section “Acoustically-conditioned latent rollout: an initial attempt”