1 of 10

Sign-SGD with Heavy Tails and Differential Privacy

Speaker: Alexey Kravatsky, 3rd student, MIPT

​

Advisor: Savelii Chezhegov, MIPT

​

03/27/2025, MIPT

2 of 10

Meld the approaches of two papers

  • Paper 1: relatively general proofs for vanilla Sign-SGD convergence
  • Paper 2: weaker proofs, but theoretical and experimental analysis of the application to the federated learning

We want to mix it to create an algorithm to securely train LLMs on user data. However, we will start with the Mushroom dataset to test our ideas.

3 of 10

Sign-SGD: send compressed gradients

Convex binary logistic regression 𝜇=0.01, 20 workers. 0.5 ms per sent bit

q(.) is a 1-bit compressor

For example, q(g) = sign(g)

4 of 10

5 of 10

6 of 10

7 of 10

The idea for our proofs is to be borrowed from Kornilov et al.

8 of 10

From China with love

9 of 10

This is impractical. Our goal is to condense it.

10 of 10

Bibliography

  • Jin et al., 2020: Chinese article with compressors and guarantees of differential privacy
  • Kornilov et al., 2025: Russian preprint with proofs of high-probability convergence in the case of heavy-tailed noise
  • These articles are key to our research. Others will be used only as a reference.