Identify a Predictive Signal for S$^3$T Backbone Generalization
Identify a consistent pre-training signal that predicts which Video-LLM backbones will benefit from S$^3$T temporal self-distillation, distinguishing successful from unsuccessful runs.
References
We examined several measurable properties of the base model, but our analysis did not identify a consistent pre-training signal that separates successful from unsuccessful S$3$T runs. We therefore leave finding such a signal for future work.
— Temporal Self-Distillation: Learning Visual State Tracking in Videos Without Supervision
(2609.04203 - Venkatraman et al., 3 Sep 2026) in Section 6, Discussion, subsection “Limitations”