Theory of optimal teacher selection for council distillation

Develop a complete theory of optimal teacher selection and council composition for distilling label-free single-cell classifiers into compact EfficientNet-B0 student models.

Background

The study finds that councils of EfficientNet-B5 and EVA-02 teachers substantially improve a compact EfficientNet-B0 student, but the numerical differences among homogeneous and mixed councils are small and do not establish a universally superior composition.

The authors explicitly identify the absence of a complete theory explaining which teacher identities, architectural relationships, or council compositions are optimal. Resolving this question would guide the construction of distillation councils beyond the limited configurations evaluated in the benchmark.

References

This supports council distillation as a strong practical strategy, but it does not yet establish a complete theory of optimal teacher selection.

Pretraining and Distillation Matter More Than Architecture Family for Label-Free Single-Cell Classification  (2609.09863 - Graemer et al., 9 Sep 2026) in Section 3.3, Knowledge distillation