Close the reading-accuracy gap between tagged input and plain kana

Improve the self-distilled pronunciation-and-accent control channel so that its target-reading accuracy on unseen words matches plain-kana input on the CosyVoice 2, Irodori-TTS-500M-v3, and T5Gemma-TTS backbones, where the tagged input currently performs 0.069–0.091 lower.

Background

The paper evaluates whether a self-distilled control channel can transfer pronunciation and accent instructions to words absent from the training carriers. On the 319-item unseen-word test set, the tagged input improves reading accuracy over unedited input on all four evaluated TTS backbones.

However, plain-kana input remains more accurate for three backbones: the tagged input is 0.091 lower on CosyVoice 2, 0.072 lower on Irodori-TTS-500M-v3, and 0.069 lower on T5Gemma-TTS. The difference is statistically significant for each of these systems, whereas the tagged and kana conditions are statistically indistinguishable on Sarashina2.2-TTS. The paper explicitly identifies closing this gap as an open problem.

References

Against plain kana, it reads $0.069$--$0.091$ lower on three of the four ($p \le 0.003$), one word in eleven to fourteen, which \S\ref{sec:conclusion} names as the open problem, and is indistinguishable on Sarashina ($-0.022$, $[-0.074,+0.033]$).

Self-Distilled Pronunciation and Accent Control for Neural Text-to-Speech  (2609.17234 - Kato, 15 Sep 2026) in Section 4.2, “Does control transfer to unseen words?”; referenced again in Section 5, “Conclusion”