1 of 13

Sparse Token Transformer with Attention Back Tracking

Heejun Lee1,2, Minki Kang1,3, Youngwan Lee1,4, Sung Ju Hwang1

KAIST1, DeepAuto.ai2, AITRICS3, ETRI4

https://github.com/gmlwns2000/sttabt

https://openreview.net/forum?id=VV0hSE8AxCw

Sparse Token Transformer with Attention Back Tracking, ICLR2023

2 of 13

Background: Light-weight Transformers

Attention Pruning

Token Pruning

2

Learned Token Pruning (LTP)

BigBird, Longformer, Reformer

Sparse Token Transformer with Attention Back Tracking, ICLR2023

3 of 13

Background: Feed-forward Token Pruning

3

Sparse Token Transformer with Attention Back Tracking, ICLR2023

4 of 13

Our Idea: Attention Back Tracking Token Pruning

4

Sparse Token Transformer with Attention Back Tracking, ICLR2023

5 of 13

Method 1: Attention Back Tracking (ABT)

5

Sparse Token Transformer with Attention Back Tracking, ICLR2023

6 of 13

Method 2: Concrete Masking (Learnable ABT)

6

References

  1. Chris et. al., The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables
  2. Yarin et. al., Concrete Dropout

[1,2]

Sparse Token Transformer with Attention Back Tracking, ICLR2023

7 of 13

Experiment Results: NLP Task - GLUE

7

Sparse Token Transformer with Attention Back Tracking, ICLR2023

8 of 13

Experiment Results: CV Task - ViT/LVViT

8

Sparse Token Transformer with Attention Back Tracking, ICLR2023

9 of 13

Experiment Results: Token Selection Visualization

9

Sparse Token Transformer with Attention Back Tracking, ICLR2023

10 of 13

Experiment Results: Overhead Analysis

10

Sparse Token Transformer with Attention Back Tracking, ICLR2023

11 of 13

Experiment Results: Computational Efficiency

11

Sparse Token Transformer with Attention Back Tracking, ICLR2023

12 of 13

Conclusion

  • We proposed two novel ideas -- Attention Back-Tracking (ABT) and Concrete masking for Transformer token pruning
  • Our STTABT methods is working well both NLP and CV tasks and various Transformer architectures (BERT, ViT, and LVViT)
  • Our STTABT methods showed better accuracy-efficiency trade-off than previous methods
  • Qualitative analysis of the tokens retained by STTABT shows that STTABT preserves relatively more important tokens, compared to previous methods

12

Sparse Token Transformer with Attention Back Tracking, ICLR2023

13 of 13

Thank You!

13

Sparse Token Transformer with Attention Back Tracking, ICLR2023