Sparse Token Transformer with Attention Back Tracking
Heejun Lee1,2, Minki Kang1,3, Youngwan Lee1,4, Sung Ju Hwang1
KAIST1, DeepAuto.ai2, AITRICS3, ETRI4
Sparse Token Transformer with Attention Back Tracking, ICLR2023
Background: Light-weight Transformers
Attention Pruning
Token Pruning
2
Learned Token Pruning (LTP)
BigBird, Longformer, Reformer
Sparse Token Transformer with Attention Back Tracking, ICLR2023
Background: Feed-forward Token Pruning
3
Sparse Token Transformer with Attention Back Tracking, ICLR2023
Our Idea: Attention Back Tracking Token Pruning
4
Sparse Token Transformer with Attention Back Tracking, ICLR2023
Method 1: Attention Back Tracking (ABT)
5
Sparse Token Transformer with Attention Back Tracking, ICLR2023
Method 2: Concrete Masking (Learnable ABT)
6
References
[1,2]
Sparse Token Transformer with Attention Back Tracking, ICLR2023
Experiment Results: NLP Task - GLUE
7
Sparse Token Transformer with Attention Back Tracking, ICLR2023
Experiment Results: CV Task - ViT/LVViT
8
Sparse Token Transformer with Attention Back Tracking, ICLR2023
Experiment Results: Token Selection Visualization
9
Sparse Token Transformer with Attention Back Tracking, ICLR2023
Experiment Results: Overhead Analysis
10
Sparse Token Transformer with Attention Back Tracking, ICLR2023
Experiment Results: Computational Efficiency
11
Sparse Token Transformer with Attention Back Tracking, ICLR2023
Conclusion
12
Sparse Token Transformer with Attention Back Tracking, ICLR2023
Thank You!
13
Sparse Token Transformer with Attention Back Tracking, ICLR2023