Sign-SGD with Heavy Tails and Differential Privacy
Meld the approaches of two papers
We want to mix it to create an algorithm to securely train LLMs on user data. However, we will start with the Mushroom dataset to test our ideas.
Sign-SGD: send compressed gradients
Convex binary logistic regression 𝜇=0.01, 20 workers. 0.5 ms per sent bit
q(.) is a 1-bit compressor
For example, q(g) = sign(g)
The idea for our proofs is to be borrowed from Kornilov et al.
From China with love
This is impractical. Our goal is to condense it.
Bibliography