Determine the added value of articulatory imaging over acoustics for contrast-specific speech changes

Determine how much of the contrast-specific loss and recovery identified from real-time MRI after glossectomy can also be recovered from acoustic signals by directly comparing audio-only and real-time-MRI-based classifiers on the same speech segments.

Background

The glossectomy analysis uses a frozen Arti-JEPA representation and phoneme probes to examine which articulatory distinctions remain decodable before and after surgery. The authors note that real-time MRI may reveal compensatory vocal-tract configurations that produce similar acoustic outcomes, making it unclear whether imaging provides information beyond the speech signal itself.

They propose comparing models trained on synchronized audio and real-time MRI under matched conditions. This would establish whether the contrast-specific changes observed in the imaging representation are also detectable acoustically or whether articulatory imaging contributes unique information for clinical assessment.

References

The key question is how much of the contrast-specific loss and recovery identified from rtMRI can also be recovered from the acoustic signal. Some articulatory distinctions may be acoustically recoverable, whereas others may remain nearly ambiguous in audio despite corresponding to distinct vocal-tract configurations; for example, canonical and compensatory articulations could produce similar acoustic outcomes while relying on different articulatory strategies. Comparing models trained on synchronized audio and rtMRI would therefore establish what additional information, if any, articulatory imaging contributes beyond acoustics alone. We leave this comparison to future work.

Arti-JEPA: Adapting Video World Model to Real-Time MRI of the Vocal Tract for Speech-Production Analysis  (2609.09757 - Nguyen et al., 9 Sep 2026) in Section 4, “Discussion on pre and post surgery monitoring and the cost of modality”