Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
97 tokens/sec
GPT-4o
53 tokens/sec
Gemini 2.5 Pro Pro
44 tokens/sec
o3 Pro
5 tokens/sec
GPT-4.1 Pro
47 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Tree-structured Auxiliary Online Knowledge Distillation (2208.10068v1)

Published 22 Aug 2022 in cs.NI

Abstract: Traditional knowledge distillation adopts a two-stage training process in which a teacher model is pre-trained and then transfers the knowledge to a compact student model. To overcome the limitation, online knowledge distillation is proposed to perform one-stage distillation when the teacher is unavailable. Recent researches on online knowledge distillation mainly focus on the design of the distillation objective, including attention or gate mechanism. Instead, in this work, we focus on the design of the global architecture and propose Tree-Structured Auxiliary online knowledge distillation (TSA), which adds more parallel peers for layers close to the output hierarchically to strengthen the effect of knowledge distillation. Different branches construct different views of the inputs, which can be the source of the knowledge. The hierarchical structure implies that the knowledge transfers from general to task-specific with the growth of the layers. Extensive experiments on 3 computer vision and 4 natural language processing datasets show that our method achieves state-of-the-art performance without bells and whistles. To the best of our knowledge, we are the first to demonstrate the effectiveness of online knowledge distillation for machine translation tasks.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (4)
  1. Wenye Lin (5 papers)
  2. Yangning Li (49 papers)
  3. Yifeng Ding (22 papers)
  4. Hai-Tao Zheng (94 papers)
Citations (1)

Summary

We haven't generated a summary for this paper yet.