1 of 54

1

Part III:

Backdoor Defenses

2 of 54

2

Pre-training stage

Training stage

backward

forward

啧啧啧啧啧啧啧啧啧在

Post-training stage

Inference stage

Attack

Defense

Secure training

Backdoor detection/mitigation

Data detection/purification

Data detection

Backdoor defense procedure

Backdoor Defense Procedure

3 of 54

3

Backdoor Defense at Different Stages

4 of 54

4

Pre-training stage

A. Backdoor Defense at Pre-training Stage

Poisoned dataset

Detection

Feature space stability

Weight space stability

Feature anomaly detection

Benign

Poisoned

5 of 54

5

A. Backdoor Defense at Pre-training Stage

6 of 54

6

Key Intuition

Activation Clustering

Detecting Backdoor Attacks on Deep Neural Networks by Activation Clustering, AAAI Workshop 2019.

  • Difference between poisoned samples and target samples are evident in latent space

(a)

(b)

(a) Activations of the target (poisoned) class. (b) Activations of the clean (unpoisoned) class.

Steps for each class:

  • Collect Activations
  • Dimension Reduction (Independent Component Analysis)
  • Clustering (k-means, DBSCAN …)
  • Cluster Analysis (Relative size, Silhouette Score …)

Weight Space Stability

Feature Space Stability

Feature Anomaly Detection

7 of 54

7

Key Intuition

Spectral Signatures in Backdoor Attacks, NeurIPS 2018.

  • The backdoor provides a strong signal in its representation for classification.

Correlations of clean samples (blue) and poisoned examples (green).

Weight Space Stability

Feature Space Stability

Spectral Signatures

Feature Anomaly Detection

Method

8 of 54

8

  • Multimodal large models exhibit surpassing capabilities of cross-modal alignment and reasoning.
  • Commonality of poisoned samples and noisy labels:
    • Visual-Linguistic Inconsistency between visual content and corrupted labels.

Versatile Data Cleanser

Key Intuition

Target Label

Poisoned Image

airplane

Ground-truth Label

automobile

VDC: Versatile Data Cleanser for Detecting Dirty Samples via Visual-Linguistic Inconsistency, arXiv 2023.

Weight Space Stability

Feature Space Stability

Feature Anomaly Detection

9 of 54

9

Versatile Data Cleanser

VDC: Versatile Data Cleanser for Detecting Dirty Samples via Visual-Linguistic Inconsistency, arXiv 2023.

Input

Target Label:

Poisoned Image:

Visual Question Generation

airplane

Large Language Model

Please give me some visual questions to ask LLM to identify if the object of image is .

Manual Design

Describe the image briefly.

Describe the image in detail.

How would you summarize the content of the image in a few words?

Provide a brief description of the given image

Q1

Q2

Q3

Q4

Is the object in the image capable of carrying people or cargo in the air?

Does the object in the image have the ability to take off on runways?

Is the object in the image designed for flying in the air?

Does the object in the image have wings for generating lift?

Q5

Q6

Q7

Q8

airplane

General Questions

Label-specific Questions

Visual Question Answering

Multimodal Large Model

Describe the image briefly.

Visual Answer Evaluation

Large Language Model

The image features a yellow car parked on a pavement, with its hood open. The car appears to be a sports car.

Please determine if the following Caption and Label refer to the same object.

Caption: [Output1]Label: [Expected answer]

Output1:

Output8:

Multimodal Large Model

Is the object in the image designed for flying in the air?

No, the object is a car.

No. Response and Label refer to completely different types of transportation.

False

True

False

False

False

True

True

False

False

Vote-based Ensemble

Poisoned Sample

Q1

Q2

Q5

Q6

Q7

Q8

Q1

Q8

Q3

Q4

String Matching

True

(Expected answer: airplane)

(Expected Answer: yes)

[Output2][Expected answer]

Weight Space Stability

Feature Space Stability

Feature Anomaly Detection

10 of 54

10

Key Intuition

Effective Backdoor Defense by Exploiting Sensitivity of Poisoned Samples, NeurIPS 2022.

  • Poisoned samples are more sensitive to transformations than clean samples.

Poisoned and clean samples with the t-SNE visualization of their features in a backdoored model

Feature Anomaly Detection

Weight Space Stability

Feature Space Stability

Feature Consistency towards Transformations

11 of 54

11

Metric: Feature Consistency towards Transformations (FCT)

Effective Backdoor Defense by Exploiting Sensitivity of Poisoned Samples, NeurIPS 2022.

Poisoned Sample Detection

Distribution of clean and poisoned samples with respect to the FCT metric on CIFAR-10.

Feature Anomaly Detection

Weight Space Stability

Feature Space Stability

Feature Consistency towards Transformations

12 of 54

12

Feature Space Stability

Feature Anomaly Detection

Key Intuition

Confusion Training

Towards A Proactive ML Approach for Detecting Backdoor Poison Samples, USENIX Security Symposium 2023.

  • Downgrade the clean accuracy while keep the Attack Success Rate .

Weight Space Stability

Steps

  • Assign random label to clean set
  • Pick a regular batch from Poisoned Dataset and a confusion batch from Clean Set
  • Joint training with large weight on confusion batch and small weight on regular batch

13 of 54

13

B. Backdoor Defense at Training Stage

Activation anomaly

Secure centralized training

Secure

model

Poisoned dataset

Secure training

14 of 54

14

B. Backdoor Defense at Training Stage

15 of 54

15

Backdoor Defense at Training Stage

Key Intuition

Anti-backdoor learning: Training clean models on poisoned data, NeurIPS 2021.

  • Backdoor task is much easier than the clean task.
  • Poisoned samples has low losses at early epochs.

Steps for each class:

  • Data Filtering: Local Gradient Ascent (LGA)
  • Model Training: Global Gradient Ascent (GGA)

Anti-Backdoor Learning

16 of 54

16

Key Intuition

Decoupling-based Backdoor Defense

Backdoor Defense via Decoupling the Training Process, ICLR 2022.

  • Poisoned samples (denoted by ‘black-cross’) tend to cluster together to form a separate cluster after the standard supervised training process,

The t-SNE of poisoned samples in the hidden space generated by Supervised Learning (a-b).

Backdoor Defense at Training Stage

17 of 54

17

Key Intuition

Decoupling-based Backdoor Defense

Backdoor Defense via Decoupling the Training Process, ICLR 2022.

  • Poisoned samples lie closely to samples with their ground-truth label after the self-supervised training process on the unlabelled poisoned dataset.

The t-SNE of poisoned samples in the hidden space generated by Self-Supervised Learning (c-d).

Backdoor Defense at Training Stage

18 of 54

18

Decoupling-based Backdoor Defense

Backdoor Defense via Decoupling the Training Process, ICLR 2022.

  1. Train the whole DNN model via self-supervised learning on label-removed training samples.
  2. Freeze the learned feature extractor and train the fully connected layers via supervised learning and filter high-credible samples based on the training loss.
  3. Adopt high-credible samples as labeled samples and remove the labels of low-credible samples to fine-tune the whole model via semi-supervised learning.

Backdoor Defense at Training Stage

19 of 54

19

Non-Adversarial Backdoor

Beating Backdoor Attack at Its Own Game, ICCV 2023.

Backdoor Defense at Training Stage

Key Intuition: Suppress the attacker's backdoor by Non-Adversarial Backdoor

  1. Preprocessing Step:
    1. Detect poisoned samples
    2. Add trigger to detected poisoned samples
    3. Relabel the detected poisoned samples by pseudo-labels
  2. Training Step: Standard Training
  3. Inference Step:
    • Add trigger to input samples and make prediction

Non-Adversarial �Trigger (Defender)�

Adversarial Trigger� (Attacker)

20 of 54

20

Server

B. Backdoor Defense at Training Stage

Secure decentralized training

Secure

model

Clients

Wight-wise defense

Malicious client detection

Aggregation

 

21 of 54

21

Server

 

 

 

 

Benign

Benign

Adversarial

X

 

Backdoor Defense at Training Stage

22 of 54

22

Weight-wise defense

Key Intuition

Pre-aggregation and Similarity Measurement

Defense against backdoor attack in federated learning, Computers & Security 121 (2022).

  • high similarity between malicious local update and global update

Method:

Malicious Client Detection

23 of 54

23

Server

 

 

 

 

 

CLIP

Gaussian

noise

Backdoor Defense at Training Stage

24 of 54

24

Method:

FLAME

FLAME: Taming Backdoors in Federated Learning, USENIX Security 2022.

  1. Filtering out backdoored models with large angular deviations
  2. Limiting the impact of scaled-up backdoors
  3. Selecting suitable noise level for backdoor elimination

Weight-wise defense

Malicious Client Detection

25 of 54

25

Server

 

 

 

 

 

 

 

 

 

 

 

Learning rate

Backdoor Defense at Training Stage

26 of 54

26

Malicious Client Detection

Weight-wise defense

Key Intuition

Robust learning rate

Defending against Backdoors in Federated Learning with Robust Learning Rate, AAAI 2021

  • Backdoor update and benign update likely differ in the directions they specify at least for some dimensions.

Method:

  • Robust learning rate: vote in form of signs of the updates to decide the learning rate

27 of 54

27

Activation anomaly

Backdoor detection

Backdoor mitigation

Post-training stage

Malicious

Secure

model

C. Backdoor Defense at Post-training Stage

28 of 54

28

Backdoor detection

Malicious

Feature-based

Weight-based

Binary classification based

Activation anomaly

C. Backdoor Defense at Post-training Stage

Backdoor detection

29 of 54

29

C. Backdoor Defense at Post-training Stage

30 of 54

30

Feature-based

Weight-based

Binary classification based

Neural cleanse: Identifying and mitigating backdoor attacks in neural networks, IEEE S&P 2019.

Neural Cleanse

  • Targeted universal adversarial perturbation (TUAP): a common perturbation that misclassifies all samples of non-targeted classes into the target class.

Large TUAP

Small TUAP on target class

Backdoor triggers create“shortcuts”from regions of a label into the region to the target(A).

Method

  • Compute TUAP for each class.
  • Detect backdoor by finding if there is one class that has much more small TUAP than others.

31 of 54

31

AEVA: Adversarial Extreme Value Analysis

AEVA: Black-Box Backdoor Detection Using Adversarial Extreme Value Analysis, ICLR 2022.

Figure: An illustration of the black-box hard-label backdoors.

Feature-based

Weight-based

Binary classification based

32 of 54

32

The distributions of perturbation values between infected (targeted) label and uninfected label are different.

AEVA: Adversarial Extreme Value Analysis

AEVA: Black-Box Backdoor Detection Using Adversarial Extreme Value Analysis, ICLR 2022.

Feature-based

Weight-based

Binary classification based

Normalized adversarial perturbation

33 of 54

33

CPBD: Critical-Path-based Backdoor Detector

Critical Path-Based Backdoor Detection for Deep Neural Networks, TNNLS 2022.

Critical Path Generation

  • Control gates G:
  • Produce Nc × K critical paths:

  • Euclidean distance (for each class)

Cross-entropy

input samples

classes

difference between the K critical paths of a specific class

Feature-based

Weight-based

Binary classification based

34 of 54

34

Universal Litmus Patterns: Revealing Backdoor Attacks in CNNs, CVPR 2020.

The general workflow of ULP on a binary property.

ULP: Universal Litmus Patterns

Step 2. ULPs and classifier h are trained.

Step 3. Backdoor detection

Feature-based

Weight-based

Binary classification based

Step 1. Train hundreds of clean models and poisoned models. Each poisoned model contain a single trigger.

Step 1. Train hundreds of clean models and poisoned models. Each poisoned model contain a single trigger.

35 of 54

35

The workflow of MNTD approach with query-tuning

MNTD: Meta Neural Trojan Detection

Detecting AI Trojans Using Meta Neural Analysis, CVPR2020.

Feature-based

Weight-based

Binary classification based

36 of 54

36

Backdoor

mitigation

Structure modification

Fine-tuning

Secure

model

Activation anomaly

C. Backdoor Defense at Post-training Stage

Backdoor removal/mitigation

37 of 54

37

C. Backdoor Defense at Post-training Stage

38 of 54

38

FP: Fine-Pruning

Fine-Pruning: Defending Against Backdooring Attacks on Deep Neural Networks, International Symposium, RAID 2018.

“Backdoor neurons” are activated when the backdoor is present in the image, while dormant in the presence of clean inputs.

Structure modification approach

Fine-tuning approach

39 of 54

39

Pruning Defense

FP: Fine-Pruning

Fine-Pruning: Defending Against Backdooring Attacks on Deep Neural Networks, International Symposium, RAID 2018.

  • Input validation clean dataset.
  • Records the average activation of each neuron.
  • Iteratively prunes Top-K neurons in increasing order of average activations.
  • Fine-tuning with clean samples.

Structure modification approach

Fine-tuning approach

40 of 54

40

Adversarial neuron perturbations

ANP: Adversarial Neuron Pruning

Adversarial Neuron Pruning Purifies Backdoored Deep Models, NeurIPS 2021.

Backdoored models are much easier to collapse and prone to output the target label than normal DNNs.

Structure modification approach

Fine-tuning approach

41 of 54

41

  • Optimize the perturbations and mask

  • Pruning neurons by their mask values

Adversarial Neuron Pruning Purifies Backdoored Deep Models, NeurIPS 2021.

Structure modification approach

Fine-tuning approach

ANP: Adversarial Neuron Pruning

42 of 54

42

CLP: Channel Lipschitzness based Pruning

Data-Free backdoor removal based on channel Lipschitzness, ECCV 2022.

  • Backdoor-related channels should have a higher Lipschitz constant than normal channels.
  • Data-free Channel Lipschitzness based Pruning

Structure modification approach

Fine-tuning approach

Channel (neuron) Lipschitzness

43 of 54

43

Pre-activation Distributions Expose Backdoor Neurons

Pre-activation Distributions Expose Backdoor Neurons, NeurIPS 2022.

Pre-activation: the maximum value in the activation map of one neuron.

Structure modification approach

Fine-tuning approach

Benign neuron: unimodal distribution of the pre-activation values

Entropy-based pruning (EP)

Given a mixed dataset, the pre-activations of each neuron are computed for each sample

Backdoor neuron: bimodal distribution of the pre-activation values

44 of 54

44

NPD: Neural Polarizer

Neural Polarizer: A Lightweight and Effective Backdoor Defense via Purifying Poisoned Feature, NeurIPS 2023.

  • Features of poisoned samples are a mixture of benign features and backdoor-related features.
  • Backdoor-related features override benign features.
  • Our goal: filtering out these backdoor-related features, while keeping benign features.

Structure modification approach

Fine-tuning approach

45 of 54

45

Three desired properties for NP

  1. Compatible with the neighboring layers.
  2. Filtering trigger features in poisoned samples.

  • Preserving benign features in poisoned and benign samples.

Approximating T and ∆

  • Dynamic approximating strategy
  • Targeted adversarial perturbation

Neural Polarizer: A Lightweight and Effective Backdoor Defense via Purifying Poisoned Feature, NeurIPS 2023.

Structure modification approach

Fine-tuning approach

NPD: Neural Polarizer

46 of 54

46

Investigating the Vanilla Fine-Tuning (FT)

FT-SAM: Sharpness-Aware Minimization

Enhancing Fine-Tuning based Backdoor Defense with Sharpness-Aware Minimization, ICCV 2023.

  • FT fails to remove backdoor effect.
  • The weights have remained mostly unchanged after FT.
  • Neurons associated with backdoors tend to exhibit large neuron weight norms.
  • Design a strategy that can significantly perturb the neurons with large weight norms, while neurons with small norms are slightly perturbed.

Structure modification approach

Fine-tuning approach

Backdoored model

TAC

47 of 54

47

  • Min-Max Formulation

Enhancing Fine-Tuning based Backdoor Defense with Sharpness-Aware Minimization, ICCV 2023.

where

  • Encourage larger perturbations to the neurons with larger weight norms

Structure modification approach

Fine-tuning approach

FT-SAM: Sharpness-Aware Minimization

  • Poisoned samples are separated in feature space.
  • The neurons, especially those with large weight norms, are more significantly perturbed via FT-SAM.

Weight perturbation

48 of 54

48

SAU: Shared Adversarial Unlearning

Shared Adversarial Unlearning: Backdoor Mitigation by Unlearning Shared Adversarial Examples, NeurIPS 2023.

Shared

Structure modification approach

Fine-tuning approach

 

Key Intuition: Trigger is a certain type of adversarial perturbation (TAP)

Observation: If the AP of poisoned model A is not a AP of model B, the trigger cannot take effect on Model B.

 

49 of 54

49

SAU: Shared Adversarial Unlearning

Shared Adversarial Unlearning: Backdoor Mitigation by Unlearning Shared Adversarial Examples, NeurIPS 2023.

Structure modification approach

Fine-tuning approach

The adversarial risk can be decomposed to three components

Informal Proposition:

Sub-Adversarial Risk:

50 of 54

50

Activation anomaly

Inference stage

Query data

repairable

No

Detect

relabelling

reject

Correct label

Yes

D. Backdoor Defense at Inference Stage

51 of 54

51

D. Backdoor Defense at Inference Stage

52 of 54

52

FreqDetector: A Frequency Perspective

Rethinking the Backdoor Attacks’ Triggers: A Frequency Perspective, ICCV 2021.

Examining images with triggers in the frequency domain using Discrete Cosine Transform

Finding: Images patched with different triggers all contain strong high-frequency components.

Method

  • Create a training set that contains clean samples and simulated poisoned data.
  • Train a binary classifier based on DCT features of the mixed training set.

Backdoor detection

Backdoor purification

Top-left: low frequencies

Right bottom : higher frequencies.

53 of 54

53

Key intuition: if we scale the input sample

  • For benign sample, the output confidence has a huge drop.
  • For poisoned sample, the output confidence is very stable.

Where and denotes a defender-specified scaling set.

Scaled prediction consistency

Scale-up: an efficient black-box input-level backdoor detection via analyzing scaled prediction consistency, ICLR 2023.

Scale-up: Scaled Prediction Consistency

Backdoor detection

Backdoor purification

54 of 54

54

Orion: Online Backdoor Sample Detection via Evolution Deviance

Orion: Online backdoor sample detection via evolution deviance. IJCAI 2023.

Deviations between the shallow and deep layers’ activations are different between benign samples and poisoned samples.

Intuition

Backdoor detection

Backdoor relabelling

Given the backdoored model, train multiple additional

early-exit classifiers.

Detection

  • For clean models: consistent and stable prediction.
  • For poisoned models: inconsistent and unstable prediction.

Relabelling: the most frequent predicted labels except for target label of these classifiers.

Method