Isolate the benefit of hard span masking in SpanDMD
Determine whether hard temporal span masks in SpanDMD independently improve or are necessary for training causal video generators to execute short-lived, stage-specific prompt conditions, as opposed to the combined effects of separate per-condition score evaluations and span-specific residual assignment.
References
A controlled comparison would use the same initialization, student schedule, and separate per-condition score evaluations, with fake-score training covering all frames whose predictions contribute to the generator update. We have not evaluated this alternative. Accordingly, our merged-prompt ablation and color-sequence diagnostic support the combined conditioning and temporal-assignment scheme of SpanDMD, but do not establish the isolated benefit or necessity of hard masking.