FT-PrivacyScore�Privacy Scoring for Machine Learning Participation
Ethan Gu, Jiajie He, Keke Chen
(Extended presentation based on a demo at ACM CCS 2024)
Outline
Human data contributors in the machine learning loop
Data collecting/
integration
Data curator/Model builder
Data contributor
Modeling
Model Deployed
Model users
Controlled access scenario
Known privacy issues with machine learning models
Data contributors’ privacy rights
https://dataprivacymanager.net/ccpa-vs-gdpr/
- Right to know potential
privacy risks!
- Right to object to processing
How do contributors learn or control privacy risks?
What existing techniques can do (with different threat model assumptions)
What we can do with the controlled access setting
Data samples are not equally sensitive!
Contributions of FT-PrivacyScore
Background: definition of privacy risk
in the context of machine learning and a sample x
Prob (model used x)
Prob (model did not use x)
Requirements for our privacy scoring method
Privacy risk estimation with membership inference (MIA)
Determine which one is more likely via hypothesis testing
Intuition: the more accurate the MIA attack, the higher the privacy risk
Membership inference method – how it works
Targeted sample x
(provided by some contributor)
Training with random samples of dataset and x
Training with random samples without x
Shadow modeling stage: repeat the above
for several times!
Purpose:
with or without x cases
Collecting enough IN-training and OUT-training examples
Learn a classification model
Or a decision rule
Out
In
What LiRA does
is the prediction confidence level at the label y
(the modeling task is classification)
Consider this is just one way to extract a feature of IN/OUT samples! There are certainly more.
p
Bird
Airplane
Truck
Dog
label y = airplane
What LiRA does
1. Collecting samples to estimate
the out-training phi distribution (blue) and
the in-training phi distribution (red);
both are approximately normal distributions
2. For a new sample, calculate its phi level’s
likelihood ratio via hypothesis testing
confobs
Benefits and challenges of LiRA
ImageNet data: 32 IN models (for red) and 32 OUT models (for blue) for determining the MIA risk for ONE sample
Offline LiRA to reduce cost (with slightly worse accuracy)
1. Train OUT models only
with public domain data (not specific to the tested sample x)
2. one-side hypothesis testing
Privacy Scoring using LiRA
Question: How well an IN sample x is correctly identified (OUT cases are not interesting)
* sampling procedure: each sample x is selected with prob. 0.5
* then for each sample x, we have something like (model1, x IN), (model2, x OUT)…
was first proposed by the “privacy onion effect” paper (NeurIPS 22)
Challenges with the privacy scoring
Our contribution – FT-PrivacyScore
A scoring service
The initial result is promising
Some promising results
Much more efficient!
Methods | Time per sample (hours) |
FT-PrivacyScore | 0.053 |
Expensive Scoring | 6.47 |
Ongoing work
Summary