1 of 16

Towards Interpretable Adversarial Examples via Sparse Adversarial Attack

1 University of Delaware,

2 Stevens Institute of Technology,

3 University of West Florida

Fudong Lin1, Jiadong Lou1, Hao Wang2, Brian Jalaian3, Xu Yuan1

2 of 16

2

Impressive Capabilities of AI Models

Natural Language Processing

Computer Vision

Robotics

Healthcare

3 of 16

3

Adversarial Attacks

  • Adversarial Attacks: Add a small and human-imperceptible noise can mislead AI models into incorrect predictions with a high confidence

Vulnerability of AI Models to Adversarial Attacks

 

Clean Image

Adversarial Example

Malicious Noise

4 of 16

4

Motivation

  • Objective: Interpret how adversarial examples mislead deep neural networks (DNNs)
  • Dense Attacks: Difficult to interpret adversarial attacks due to their overly perturbed adversarial examples
  • Sparse Attacks: NP-hard problem, with their resultant adversarial examples suffering from poor sparsity

Interpretable Adversarial Attacks

Our Goal: Develop a sparse attack that yields interpretable adversarial examples, allowing us to interpret the vulnerability of DNNs

5 of 16

5

Our Sparse Attack

 

 

 

  • Loss Function: Maximize the classification error and minimize the number of perturbed pixels simultaneously:

Intractable!

6 of 16

6

Our Sparse Attack

  • Reparameterization: Reformulate the loss function by using the Heaviside step function:

 

7 of 16

7

Our Sparse Attack

  • Dirac Delta Function: The distributional derivative of the Heaviside step function—this function is zero everywhere except at the origin, which is infinite:
  • Zero-Centered Normal Distribution:

 

Non-differentiable!

8 of 16

8

Our Sparse Attack

 

9 of 16

9

Evaluation on the Sparsity

Targeted Attack

Our approach yields adversarial examples with the best sparsity on both scenarios.

 

Non-Targeted Attack

10 of 16

10

Interpret the Vulnerability of DNNs

Our approach can interpret the vulnerability of various DNN architectures.

  • Explainable AI: Grad-CAM and Guided Grad-CAM
  • Two Types of Malicious Noises: “obscuring noise”, which prevents DNNs from identify true labels; and “leading noise”, which leads DNNs into incorrect decisions.

ResNet-50: Mislead “canoe” to “wings”

VGG-16: Mislead “candle” to “toilet tissue”

11 of 16

11

Attack Robust Models

 

Comparison to Sparse Attack Counterparts

Our approach achieves the best attack successful rates under all scenarios.

12 of 16

12

Evaluation on the Attack Transferability

 

Comparison of Transferability under ImageNet

Our approach achieves the best transferability under all scenarios.

13 of 16

13

Comparison on Computation Complexity

 

Comparison of Time Cost under ImageNet

Our approach runs much faster than all sparse attack counterparts.

14 of 16

14

Ablation Studies

 

Hyperparameter Sensitivity

Finetuning hyperparameters yields better trade-off between fooling rate and sparsity.

15 of 16

15

Conclusion

    • This work has proposed a sparse attack that yields interpretable adversarial examples, thanks to their superb sparsity.
    • Our key idea is to approximate the NP-hard sparsity optimization problem via a theoretical sound reparameterization technique. This makes direct optimization of sparse perturbations computationally tractable.
    • Our approach outperforms sparse attack counterparts in terms of computation efficiency, transferability, and attack intensity.
    • Our approach helps reveal two types of adversarial perturbations. This empowers us to interpret how adversarial examples mislead DNNs into incorrect decisions.

Code

Paper

16 of 16

16

Thank you!

Q & A