Differential Privacy �as a privacy-preserving �technique in �Federated Learning
Final Presentation
Master Thesis No. : MT- 3506
Institute of Industrial Automation
and Software Engineering
Karthik Rajendran
Master in Electrical Engineering
Motivation
Distributed Machine Learning
Clients send data to a central server for model training, and the server returns the trained model to clients.
Clients share their sensitive data which raises concerns about data privacy.
Federated Learning
Models are trained locally with local client data. Server aggregates local weights and returns global weights.
Regeneration of local data via model inversion from external threats.
Adding noise via Differential Privacy to protect local weights.
University of Stuttgart, IAS
2
10/4/2023
How to train ML models in a distributed manner while keeping sensitive information safe from external threats ?
Early Pioneers of Federated Learning
University of Stuttgart, IAS
3
10/4/2023
Source: Huili Chen, Amazon.science
Source: Brendan McMahan and Daniel Ramage, Research Scientists, Google Research
Amazon uses Federated Learning to improve their smart devices.
Google uses Federated Learning to improve keyboard suggestions.
Contents
University of Stuttgart, IAS
4
10/4/2023
Problem Statement
Developed Concept
Implementation
Results and Evaluation
Summary and Future Scope
Federated Learning
University of Stuttgart, IAS
5
10/4/2023
Central Server
Client1
Client2
Client3
Federated Learning
University of Stuttgart, IAS
6
10/4/2023
Client1
Client2
Client3
Central Server
Federated Learning
University of Stuttgart, IAS
7
10/4/2023
Client1
Client2
Client3
Central Server
Federated Learning
University of Stuttgart, IAS
8
10/4/2023
Client1
Client2
Client3
Central Server
Federated Learning
University of Stuttgart, IAS
9
10/4/2023
Clients avoid sharing their training data which improves privacy conditions.
Client1
Client2
Client3
Client1
Central Server
Security Threat
University of Stuttgart, IAS
10
10/4/2023
GAN
Hacker gets the local weight
Client local weights
Model Inversion Attack
Client1
Client2
Client3
Central Server
Differential Privacy
Safeguards individual privacy by ensuring that specific information about individuals cannot be deduced.
University of Stuttgart, IAS
11
10/4/2023
Outcome is approximately the same
Differential Privacy Integration
University of Stuttgart, IAS
12
10/4/2023
Data Collection
Data Preprocessing
Data Labelling
Model Training
Model evaluation and Testing
Data Sharing and Collaboration
Querying the Model
Model Deployment
Data Reporting and Summarization
Post Processing
Data
Model
Inference
Differentially Private Stochastic Gradient Descent (DPSGD)
University of Stuttgart, IAS
13
10/4/2023
Traditional Model Training
Differentially Private Model Training
Find the compromise between privacy and model accuracy
University of Stuttgart, IAS
14
10/4/2023
„
Contents
University of Stuttgart, IAS
15
10/4/2023
Problem Statement
Developed Concept
Implementation
Results and Evaluation
Summary and Future Scope
System Architecture
University of Stuttgart, IAS
16
10/4/2023
Model Training
Server
Driver
2
Sensors
1
1
Driver Behaviour Data
2
Client System
3
Local weights
4
Global weights
Data Protection
3
4
. . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . .
5
5
Server
Driver Behaviour Data
University of Stuttgart, IAS
17
10/4/2023
1
Model Training
University of Stuttgart, IAS
18
10/4/2023
2
Server
Multiple clients collaboratively train a global model by sharing model updates while preserving data privacy on the individual clients using Federated Average.
University of Stuttgart, IAS
19
10/4/2023
w1
w2
w3
5
Central Server
Local weights
Global
weights
Gw
Gw
System process
University of Stuttgart, IAS
20
10/4/2023
1
2
3
4
5
Initialization
Initialization of a global model in the server
Client Selection
Subset of clients is selected to participate
Local Private Model Training
Clients train on the global model
Model Aggregation
Server aggregates the client weights
Model Update Distribution
New global model is distributed back
6
7
8
9
10
Iterations
Repeated for multiple rounds
Convergence
Global model converges to high accuracy
Stopping Criteria
Training rounds continue until criteria is met
Final Global Model
Global model can be used for inference
Model Evaluation
Final global model is evaluated
Contents
University of Stuttgart, IAS
21
10/4/2023
Problem Statement
Developed Concept
Implementation
Results and Evaluation
Summary and Future Scope
Dataset
Driving Dataset from OCSlab, Korea University
University of Stuttgart, IAS
22
10/4/2023
Intake air pressure | Fuel consumption | Engine torque | Time (s) | . . . . . . | Driver |
33 | 268.8 | 5.5 | 1 | | 1 |
40 | 243.2 | 7 | 2 | | 1 |
41 | 217.6 | 7 | 3 | | 1 |
Dataset Preprocessing
University of Stuttgart, IAS
23
10/4/2023
Architecture Implementation
Raspberry Pi devices are utilized as automotive clients, conducting model training
and testing within a Dockerized container utilizing the Flower framework.
University of Stuttgart, IAS
24
10/4/2023
Central Server
Local weights
Global weights
Architecture Components
University of Stuttgart, IAS
25
10/4/2023
Model
DPSGD Optimizer
Privacy Accountant
Architecture Components
University of Stuttgart, IAS
26
10/4/2023
Model
DPSGD Optimizer
Privacy Accountant
*DPSGD (Differentially Private Stochastic Gradient Descent)
Gradient
Gradient Clipping
Noise Addition
Clipping Bound
DPSGD Gradient
Architecture Components
University of Stuttgart, IAS
27
10/4/2023
Model
DPSGD Optimizer
Privacy Accountant
Privacy Utility Optimization
University of Stuttgart, IAS
28
10/4/2023
Noise
Privacy
Accuracy
Optimal value
Select optimal privacy with slight accuracy loss.
Experiments
Training and Testing
University of Stuttgart, IAS
29
10/4/2023
Timesteps | 100 |
Rounds | 100 |
Batch Size | 32 |
Optimizer | DPSGD |
Learning rate | 0.001 |
Clipping Bound | Noise Multiplier |
5 | 0.1 |
5 | 0.5 |
5 | 0.8 |
5 | 1 |
5 | 5 |
5 | 10 |
Contents
University of Stuttgart, IAS
30
10/4/2023
Problem Statement
Developed Concept
Implementation
Results and Evaluation
Summary and Future Scope
Accuracy Across Training Rounds and Noise Levels
University of Stuttgart, IAS
31
10/4/2023
🡪 Increasing noise, reduces accuracy
Clipping Bound = 5
Accuracy, Privacy Budget and Noise Comparison
University of Stuttgart, IAS
32
10/4/2023
Client 1
Client 2
Client 3
🡪 As the noise increases, the accuracy reduces and total privacy improves.
Confusion Matrix on Validation Data
University of Stuttgart, IAS
33
10/4/2023
Noise = 0.1
Noise = 10
🡪 Model prediction across class reduces as noise increases.
Correlation Matrix of Train Accuracy
University of Stuttgart, IAS
34
10/4/2023
Noise | Client Averaged Train Accuracy | Privacy Accountant |
0.8 | 65% | 774 |
1.0 | 63% | 523 |
Selecting a noise level of 1.0 over 0.8 ensures a 32% increase in model privacy with minimal impact on accuracy.
Noise
Noise
Contents
University of Stuttgart, IAS
35
10/4/2023
Problem Statement
Developed Concept
Implementation
Results and Evaluation
Summary and Future Scope
Summary and Future Scope
University of Stuttgart, IAS
36
10/4/2023
References
[1] ocslab.hksecurity.net/Datasets/driving-dataset
University of Stuttgart, IAS
37
10/4/2023
Karthik Rajendran
Institute for Industrial Automation and Software Engineering
Prof. Dr.-Ing. Dr. h.c. Michael Weyrich
st176764@stud.uni-stuttgart.de
University of Stuttgart
Thank you!