Isolate the benefit of hard span masking in SpanDMD

Determine whether hard temporal span masks in SpanDMD independently improve or are necessary for training causal video generators to execute short-lived, stage-specific prompt conditions, as opposed to the combined effects of separate per-condition score evaluations and span-specific residual assignment.

Background

SpanDMD evaluates each active prompt on the complete noised rollout while retaining the corresponding DMD residual only within that prompt’s temporal span. The paper motivates this masking strategy by arguing that applying every condition’s residual to all frames could cause competing supervision from prompts describing incompatible stages of a transition.

However, the reported ablation removes SpanDMD by merging all conditions into one teacher prompt and therefore changes both score-network conditioning and temporal residual assignment. The paper explicitly notes that this experiment cannot isolate the contribution or necessity of hard masking itself. A controlled comparison would retain separate per-condition score evaluations and the same student schedule while removing only the span masks.

References

A controlled comparison would use the same initialization, student schedule, and separate per-condition score evaluations, with fake-score training covering all frames whose predictions contribute to the generator update. We have not evaluated this alternative. Accordingly, our merged-prompt ablation and color-sequence diagnostic support the combined conditioning and temporal-assignment scheme of SpanDMD, but do not establish the isolated benefit or necessity of hard masking.

— No Corners Cut: State-Grounded Transitions for Mid-Stream Prompt Switches in Video Generation  (2609.38691 - Rao et al., 30 Sep 2026) in Appendix A.6, subsection “Why use span masks?”