Data Attribution for Text-to-Image Models by Unlearning Synthesized Images
Richard Zhang
Jun-Yan Zhu
Alexei A. Efros
Sheng-Yu Wang
In NeurIPS, 2024.
Aaron Hertzmann
Influence scores
GenAI image
Data Attribution
Dataset
Stable Diffusion
Challenge: ground truth influence is unknown…
Must intervene in the training process
Random subsets
Training�dataset
Counterfactual subset 1
Analyze models
Counterfactual subset 2N
Training 2N models is too expensive
c.f. Feldman & Zhang. What Neural Networks Memorize and Why. NeurIPS 2020.
Synthesized
Leave-one-out
Training�dataset
Unlearning
Koh & Liang. ICML 2017; Schioppa AAAI 2022; Park ICML 2023; Georgiev ICML Wkshp 2023; Grosse ArXiv 2023.
Evaluate
Influence functions: linear approx.�for unlearning & evaluation
Synthesized
Store a low-dimensional version�or recompute at test-time
Leave-one-out
Training�dataset
Synthesized
Can we remove the linear approximations?
Storing or recomputing�N models intractable
Unlearning
Evaluate
Koh & Liang. ICML 2017; Schioppa AAAI 2022; Park ICML 2023; Georgiev ICML Wkshp 2023; Grosse ArXiv 2023.
Attribution by Unlearning (AbU)
Assess influence
(by loss increase)
Training�dataset
Unlearning
Synthesized
Counterfactual evaluation
Training�dataset
Remove Top-K�Influential Images
Counterfactual subset
c.f. K. Georgiev, et al. How Training Data Guides Diffusion Models. In ArXiv, 2023.
Synthesized
Influences
Counterfactual evaluation
Failed Resynthesis
Training�dataset
Remove Top-K�Influential Images
Counterfactual subset
2. Regenerate & Assess difference
1. Check DDPM loss
If critical training images are identified,�removing them should destroy the generation
Expensive evaluation…
…but let’s do it! (for modest sizes)
c.f. K. Georgiev, et al. How Training Data Guides Diffusion Models. In ArXiv, 2023.
Synthesized
Influences
MS-COCO results
Ours
“A bus traveling on a freeway next to other traffic.”
DINO
JourneyTRAK
c.f. K. Georgiev, et al. How Training Data Guides Diffusion Models. In ArXiv, 2023.
Remove K=500
(0.4% of dataset)
Ours
DINO
JourneyTRAK
Attribution results
D-TRAK
D-TRAK
c.f. X. Zheng, et al. Intriguing Properties of Data Attribution on Diffusion Models. In ICLR, 2024.
Effective removal
Nearest neighbors
Tuned via Customization
Influence functions
Local attribution
“A motorcycle and a stop sign.”
Cropped
Queries
Attributed training images
Customized Model Benchmark
Object Centric
Generated Sample
DINO (AbC)
Ours
CLIP (AbC)
“An ink drawing of�V* turtle”
“A picture of tree in the style of V* art”
Ours
DINO (AbC)
CLIP (AbC)
Ours
D-TRAK
D-TRAK
Thank you!