Frame-level versus phone-level CDE evaluation

Determine whether a neural controlled differential equation operating at one frame per solver step provides a richer temporal progression than a neural controlled differential equation operating at one phone per solver step, while accounting for the resulting increase in ordinary differential equation solver steps.

Background

The proposed text-to-speech system uses a neural controlled differential equation driven by a duration-augmented phone-level control path. In the reported experiments, the CDE is evaluated using solver step sizes corresponding to phone-level processing, including one phone per step and half-phone steps.

Because continuous-time models can, in principle, be evaluated at finer temporal resolutions, the paper identifies frame-level evaluation as a possible alternative. Such evaluation could provide a more detailed temporal progression within each phone, but it would require substantially more ordinary differential equation solver steps and therefore incur additional computational cost. The paper leaves the relative modelling benefit of this finer resolution unresolved.

References

Although the reparametrisation invariance theorem suggests operating in label-space provides a sufficiently strong model, an interesting question left unanswered in this work is whether a CDE operating one frame per step rather than one phone per step provides a richer temporal progression at the cost of more ODE solver steps.

Continuous-Time Acoustic Modelling with Neural Controlled Differential Equations  (2609.11725 - Cross et al., 10 Sep 2026) in Section 3, subsection “Implementation Details”