1 of 14

Guiding the Student’s Learning Curve: Augmenting

Knowledge Distillation with Insights from

GradCAM

Suvaditya Mukherjee

Department of Artificial Intelligence

NMIMS University, Mumbai

suvaditya.mukherjee015@nmims.edu.in

Dev Chandan

Department of Artificial Intelligence

NMIMS University, Mumbai

dev.chandan027@nmims.edu.in

Shreyas Dongre

Department of Artificial Intelligence

NMIMS University, Mumbai

shreyas.dongre134@nmims.edu.in

2 of 14

2

Roadmap

  • Introduction
  • Related works
  • Proposed Methodology
    • GradCAM
    • Distillation
    • Information Fusion
    • Model Inference
  • Experiments
  • Conclusion

3 of 14

  • Our proposed technique involves leveraging GradCAM outputs

as an additional input to the Student network for improved representation learning.

  • Our findings reveal that this approach facilitates

expedited convergence, particularly when the Teacher network

exhibits strong performance and a substantial size advantage

over the Student network.

Introduction

4 of 14

Ref

Authors

Title

Key Contributions

Year

[1]

Ramprasaath R. Selvaraju, Michael Cogswell et. al.

Grad-cam: Visual

explanations from deep networks via gradient-based localization

A model explainability technique that allows us to understand underlying representations learnt by a Convolutional Layer that has learnable kernels

2020

[2]

Geoffrey Hinton, Oriol Vinyals, and Jeff Dean

Distilling the knowledge

in a neural network

Introduces a technique that allows us to transfer learnings from a larger network to a smaller one

2015

[3]

Jangho Kim, Yash Bhalgat, Jinwon Lee, Chirag Patel, et. al.

Quantization-aware knowledge distillation

Allows us to extend distillation with quantization-aware training

2019

[4]

Ding Zeyu, Razali Yaakob, Azreen Azman et. al.

A grad-cam-

based knowledge distillation method for the detection of tuberculosis

A method that allows us to introduce a new loss term that compares the representations of the teacher network and the student network.

2022

Literature Survey

5 of 14

Proposed Methodology

6 of 14

6

GradCAM

7 of 14

7

Distillation

8 of 14

Information Fusion

9 of 14

9

Model Inference

10 of 14

Experiments

Training on CIFAR-10

11 of 14

Experiments

Training on CIFAR-10

12 of 14

Experiments

Training on CIFAR-10

13 of 14

Conclusion

By introducing GradCAM as a secondary input for the

Student network to learn from, we’ve taken a significant leap

from traditional distillation process and have opened up an

additional backdoor to enhance the distillation process, thereby

enabling students to converge efficiently.

14 of 14

Thank You