Establish whether tuned feature-based distillation transfers representation-space unknown-detection ability

Establish whether feature-based distillation with a tuned loss weight, evaluated across several training start dates, can transfer the teacher’s representation-space unknown-detection advantage to a compact encrypted-traffic classifier.

Background

The exploratory experiments locate the teacher’s unknown-detection advantage in its penultimate representation rather than in its logits. A similarity-preserving feature-distillation student, trained with one untuned loss weight and on only one start date, transfers less of that advantage than logit distillation. Because the loss weight was not tuned and the experiment was not replicated across start dates, the result does not determine whether feature-based distillation can preserve representation-space detection ability under a suitable objective and training configuration.

References

Feature-based distillation remains the most important question this study leaves open, all the more so now that Section~\ref{sec:featurescores} places the teacher's advantage in the representation and not in the logits.

— Unknown-Traffic Detection, Calibration and Shortcut Reliance in Distilled Encrypted-Traffic Classifiers over One Year  (2609.31141 - Abbasi, 25 Sep 2026) in Section 6, “Limitations”; Section 7, “Conclusion and Future Work”