Close the reading-accuracy gap between tagged input and plain kana
Improve the self-distilled pronunciation-and-accent control channel so that its target-reading accuracy on unseen words matches plain-kana input on the CosyVoice 2, Irodori-TTS-500M-v3, and T5Gemma-TTS backbones, where the tagged input currently performs 0.069–0.091 lower.
References
Against plain kana, it reads $0.069$--$0.091$ lower on three of the four ($p \le 0.003$), one word in eleven to fourteen, which \S\ref{sec:conclusion} names as the open problem, and is indistinguishable on Sarashina ($-0.022$, $[-0.074,+0.033]$).
— Self-Distilled Pronunciation and Accent Control for Neural Text-to-Speech
(2609.17234 - Kato, 15 Sep 2026) in Section 4.2, “Does control transfer to unseen words?”; referenced again in Section 5, “Conclusion”