跳到主要导航 跳到搜索 跳到主要内容

Adaptive multi-teacher multi-level knowledge distillation

  • East China Normal University

科研成果: 期刊稿件文章同行评审

摘要

Knowledge distillation (KD) is an effective learning paradigm for improving the performance of lightweight student networks by utilizing additional supervision knowledge distilled from teacher networks. Most pioneering studies either learn from only a single teacher in their distillation learning methods, neglecting the potential that a student can learn from multiple teachers simultaneously, or simply treat each teacher to be equally important, unable to reveal the different importance of teachers for specific examples. To bridge this gap, we propose a novel adaptive multi-teacher multi-level knowledge distillation learning framework (AMTML-KD), which consists two novel insights: (i) associating each teacher with a latent representation to adaptively learn instance-level teacher importance weights which are leveraged for acquiring integrated soft-targets (high-level knowledge) and (ii) enabling the intermediate-level hints (intermediate-level knowledge) to be gathered from multiple teachers by the proposed multi-group hint strategy. As such, a student model can learn multi-level knowledge from multiple teachers through AMTML-KD. Extensive results on publicly available datasets demonstrate the proposed learning framework ensures student to achieve improved performance than strong competitors.

源语言英语
页(从-至)106-113
页数8
期刊Neurocomputing
415
DOI
出版状态已出版 - 20 11月 2020

学术指纹

探究 'Adaptive multi-teacher multi-level knowledge distillation' 的科研主题。它们共同构成独一无二的学术指纹。

引用此