Towards Interpretable Adversarial Examples via Sparse Adversarial Attack
1 University of Delaware,
2 Stevens Institute of Technology,
3 University of West Florida
Fudong Lin1, Jiadong Lou1, Hao Wang2, Brian Jalaian3, Xu Yuan1
2
Impressive Capabilities of AI Models
Natural Language Processing
Computer Vision
Robotics
Healthcare
3
Adversarial Attacks
Vulnerability of AI Models to Adversarial Attacks
Clean Image
Adversarial Example
Malicious Noise
4
Motivation
Interpretable Adversarial Attacks
Our Goal: Develop a sparse attack that yields interpretable adversarial examples, allowing us to interpret the vulnerability of DNNs
5
Our Sparse Attack
Intractable!
6
Our Sparse Attack
7
Our Sparse Attack
Non-differentiable!
8
Our Sparse Attack
9
Evaluation on the Sparsity
Targeted Attack
Our approach yields adversarial examples with the best sparsity on both scenarios.
Non-Targeted Attack
10
Interpret the Vulnerability of DNNs
Our approach can interpret the vulnerability of various DNN architectures.
ResNet-50: Mislead “canoe” to “wings”
VGG-16: Mislead “candle” to “toilet tissue”
11
Attack Robust Models
Comparison to Sparse Attack Counterparts
Our approach achieves the best attack successful rates under all scenarios.
12
Evaluation on the Attack Transferability
Comparison of Transferability under ImageNet
Our approach achieves the best transferability under all scenarios.
13
Comparison on Computation Complexity
Comparison of Time Cost under ImageNet
Our approach runs much faster than all sparse attack counterparts.
14
Ablation Studies
Hyperparameter Sensitivity
Finetuning hyperparameters yields better trade-off between fooling rate and sparsity.
15
Conclusion
Code
Paper
16
Thank you!
Q & A