Adaptive Multi-Teacher Knowledge Distillation with Meta-Learning (2306.06634v1)

Published 11 Jun 2023 in cs.CV

Abstract: Multi-Teacher knowledge distillation provides students with additional supervision from multiple pre-trained teachers with diverse information sources. Most existing methods explore different weighting strategies to obtain a powerful ensemble teacher, while ignoring the student with poor learning ability may not benefit from such specialized integrated knowledge. To address this problem, we propose Adaptive Multi-teacher Knowledge Distillation with Meta-Learning (MMKD) to supervise student with appropriate knowledge from a tailored ensemble teacher. With the help of a meta-weight network, the diverse yet compatible teacher knowledge in the output layer and intermediate layers is jointly leveraged to enhance the student performance. Extensive experiments on multiple benchmark datasets validate the effectiveness and flexibility of our methods. Code is available: https://github.com/Rorozhl/MMKD.

PDF HTML Abstract

Summarize Bookmark Chat (Pro)

Authors (3)

Hailin Zhang (51 papers)
Defang Chen (28 papers)
Can Wang (156 papers)

Citations (5)

View on Semantic Scholar

GitHub

GitHub - Rorozhl/MMKD: This is the implementation for the ICME-2023 paper (Adaptive Multi-Teacher Knowledge Distillation with Meta-Learning). (20 stars)

Adaptive Multi-Teacher Knowledge Distillation with Meta-Learning (2306.06634v1)

Related Papers

GitHub