Skip to main navigation Skip to search Skip to main content

Contrastive Learning or Masked Autoencoder? Understanding and Improving Self-Supervised Knowledge Distillation

Research output: Contribution to journalArticlepeer-review

Abstract

Lying at the intersection of self-supervised learning (SSL) and knowledge distillation (KD), Self-supervised KD (SSKD) differs from classical KD frameworks by assuming the teacher model is pretrained without labels. As SSKD relies on SSL-pretrained teachers, recent works have largely employed masked image modeling (MIM) based teachers, especially masked autoencoder (MAE). In this work, however, we explore the previously untapped potential of contrastive learning (CL) as a teacher in self-supervised knowledge distillation. Our findings reveal that 1) CL-pretrained teachers outperform MIM-pretrained ones in student training due to their richer semantic representations, and 2) despite this advantage, student distilled from CL exhibit a lack of attention diversity. Remarkably, we observe that MAE, while less effective for direct distillation, exhibits richer attention patterns than CL. Motivated by this complementary property, we integrate the distillation of MAE attention scores into the CL-based SSKD framework, named I-SSKD. Our approach effectively enhances attention diversity and demonstrates improved performance in downstream visual recognition tasks, including ImageNet-1K classification, MS-COCO object detection, and ADE20K semantic segmentation. In addition, it outperforms current state-of-the-art methods, establishing a strong baseline for self-supervised knowledge distillation.

Original languageEnglish
Pages (from-to)39472-39482
Number of pages11
JournalIEEE Access
Volume14
DOIs
Publication statusPublished - 2026

Bibliographical note

Publisher Copyright:
© 2013 IEEE

Keywords

  • contrastive learning
  • Deep learning
  • knowledge distillation
  • masked image modeling
  • self-supervised learning

Fingerprint

Dive into the research topics of 'Contrastive Learning or Masked Autoencoder? Understanding and Improving Self-Supervised Knowledge Distillation'. Together they form a unique fingerprint.

Cite this