Identify a Predictive Signal for S$^3$T Backbone Generalization

Identify a consistent pre-training signal that predicts which Video-LLM backbones will benefit from S$^3$T temporal self-distillation, distinguishing successful from unsuccessful runs.

Background

S3^3T does not generalize consistently across the eleven base models from five model families tested in the paper: only LLaVA-OV-2-8B showed a clear improvement. The authors investigated ten possible explanations but found none that accounted for the difference. They identify video post-training amount and the extent to which a low-rank language-model adapter can alter visual-evidence use as remaining possibilities, but public checkpoints do not permit these factors to be tested independently.

The unresolved problem is to determine a measurable, consistent pre-training signal that separates backbones likely to benefit from S3^3T from those that will not. Solving it would support principled backbone selection and clarify why the method’s transferability varies across architectures and training histories.

References

We examined several measurable properties of the base model, but our analysis did not identify a consistent pre-training signal that separates successful from unsuccessful S$3$T runs. We therefore leave finding such a signal for future work.

Temporal Self-Distillation: Learning Visual State Tracking in Videos Without Supervision  (2609.04203 - Venkatraman et al., 3 Sep 2026) in Section 6, Discussion, subsection “Limitations”